mirror of
https://github.com/BobTheBlinker/android_kernel_motorola_sm6375.git
synced 2026-10-07 04:12:04 -04:00
Merge branch 'wireguard' into lineage-20
Changes are cherry-picked from kernel/common @ android12-5.4
into a branch called "wireguard" and then merged into lineage-20
via: `git merge --log=99999999 --no-ff wireguard`.
-----
By Jason A. Donenfeld (92) and others
* wireguard:
UPSTREAM: wireguard: send: annotate intentional data race in checking empty queue
UPSTREAM: wireguard: queueing: annotate intentional data race in cpu round robin
UPSTREAM: wireguard: allowedips: avoid unaligned 64-bit memory accesses
UPSTREAM: wireguard: netlink: access device through ctx instead of peer
UPSTREAM: wireguard: netlink: check for dangling peer via is_dead instead of empty list
UPSTREAM: wireguard: receive: annotate data-race around receiving_counter.counter
UPSTREAM: wireguard: allowedips: expand maximum node depth
UPSTREAM: wireguard: netlink: send staged packets when setting initial private key
UPSTREAM: wireguard: queueing: use saner cpu selection wrapping
UPSTREAM: wireguard: netlink: avoid variable-sized memcpy on sockaddr
UPSTREAM: wireguard: ratelimiter: disable timings test by default
UPSTREAM: wireguard: allowedips: don't corrupt stack when detecting overflow
UPSTREAM: wireguard: ratelimiter: use hrtimer in selftest
UPSTREAM: crypto: arm64/poly1305 - fix a read out-of-bound
UPSTREAM: wireguard: device: check for metadata_dst with skb_valid_dst()
UPSTREAM: wireguard: socket: ignore v6 endpoints when ipv6 is disabled
UPSTREAM: wireguard: socket: free skb in send6 when ipv6 is disabled
UPSTREAM: wireguard: queueing: use CFI-safe ptr_ring cleanup function
UPSTREAM: wireguard: allowedips: add missing __rcu annotation to satisfy sparse
UPSTREAM: wireguard: ratelimiter: use kvcalloc() instead of kvzalloc()
UPSTREAM: wireguard: receive: drop handshakes if queue lock is contended
ANDROID: properly copy the scm_io_uring field in struct sk_buff
UPSTREAM: wireguard: receive: use ring buffer for incoming handshakes
UPSTREAM: wireguard: device: reset peer src endpoint when netns exits
UPSTREAM: wireguard: selftests: actually test for routing loops
UPSTREAM: wireguard: selftests: increase default dmesg log size
UPSTREAM: crypto: x86/curve25519 - fix cpu feature checking logic in mod_exit
BACKPORT: wireguard: peer: allocate in kmem_cache
UPSTREAM: crypto: mips: add poly1305-core.S to .gitignore
UPSTREAM: crypto: poly1305 - fix poly1305_core_setkey() declaration
UPSTREAM: crypto: mips/poly1305 - enable for all MIPS processors
UPSTREAM: wireguard: kconfig: use arm chacha even with no neon
UPSTREAM: wireguard: queueing: get rid of per-peer ring buffers
UPSTREAM: wireguard: device: do not generate ICMP for non-IP packets
UPSTREAM: wireguard: selftests: test multiple parallel streams
UPSTREAM: crypto: Kconfig - CRYPTO_MANAGER_EXTRA_TESTS requires the manager
UPSTREAM: wireguard: allowedips: free empty intermediate nodes when removing single node
UPSTREAM: wireguard: allowedips: allocate nodes in kmem_cache
UPSTREAM: wireguard: selftests: remove old conntrack kconfig value
UPSTREAM: wireguard: allowedips: remove nodes in O(1)
UPSTREAM: wireguard: allowedips: initialize list head in selftest
UPSTREAM: wireguard: selftests: make sure rp_filter is disabled on vethc
UPSTREAM: wireguard: use synchronize_net rather than synchronize_rcu
UPSTREAM: wireguard: do not use -O3
UPSTREAM: crypto: arm/curve25519 - Move '.fpu' after '.arch'
UPSTREAM: crypto: jitter - SP800-90B compliance
UPSTREAM: crypto: jitter - add header to fix buildwarnings
UPSTREAM: crypto: jitter - fix comments
UPSTREAM: wireguard: Kconfig: select CRYPTO_BLAKE2S_ARM
FROMLIST: crypto: arm64/poly1305-neon - reorder PAC authentication with SP update
crypto: lib/sha256 - add sha256() function
crypto: lib/sha256 - return void
UPSTREAM: wireguard: peerlookup: take lock before checking hash in replace operation
UPSTREAM: wireguard: noise: take lock when removing handshake entry from table
UPSTREAM: netlink: consistently use NLA_POLICY_MIN_LEN()
UPSTREAM: netlink: consistently use NLA_POLICY_EXACT_LEN()
UPSTREAM: wireguard: queueing: make use of ip_tunnel_parse_protocol
UPSTREAM: wireguard: implement header_ops->parse_protocol for AF_PACKET
UPSTREAM: net: ip_tunnel: add header_ops for layer 3 devices
UPSTREAM: wireguard: receive: account for napi_gro_receive never returning GRO_DROP
UPSTREAM: wireguard: device: avoid circular netns references
UPSTREAM: wireguard: noise: do not assign initiation time in if condition
UPSTREAM: wireguard: noise: separate receive counter from send counter
UPSTREAM: wireguard: queueing: preserve flow hash across packet scrubbing
UPSTREAM: wireguard: noise: read preshared key while taking lock
UPSTREAM: wireguard: selftests: use newer iproute2 for gcc-10
UPSTREAM: wireguard: send/receive: use explicit unlikely branch instead of implicit coalescing
UPSTREAM: wireguard: selftests: initalize ipv6 members to NULL to squelch clang warning
UPSTREAM: wireguard: send/receive: cond_resched() when processing worker ringbuffers
UPSTREAM: wireguard: socket: remove errant restriction on looping to self
UPSTREAM: wireguard: selftests: use normal kernel stack size on ppc64
UPSTREAM: wireguard: receive: use tunnel helpers for decapsulating ECN markings
UPSTREAM: wireguard: queueing: cleanup ptr_ring in error path of packet_queue_init
UPSTREAM: wireguard: send: remove errant newline from packet_encrypt_worker
UPSTREAM: wireguard: noise: error out precomputed DH during handshake rather than config
UPSTREAM: wireguard: receive: remove dead code from default packet type case
UPSTREAM: wireguard: queueing: account for skb->protocol==0
UPSTREAM: wireguard: selftests: remove duplicated include <sys/types.h>
UPSTREAM: wireguard: socket: remove extra call to synchronize_net
UPSTREAM: wireguard: send: account for mtu=0 devices
UPSTREAM: wireguard: receive: reset last_under_load to zero
UPSTREAM: wireguard: selftests: reduce complexity and fix make races
UPSTREAM: wireguard: device: use icmp_ndo_send helper
UPSTREAM: wireguard: selftests: tie socket waiting to target pid
UPSTREAM: wireguard: selftests: ensure non-addition of peers with failed precomputation
UPSTREAM: wireguard: noise: reject peers with low order public keys
UPSTREAM: wireguard: allowedips: fix use-after-free in root_remove_peer_lists
UPSTREAM: wireguard: socket: mark skbs as not on list when receiving via gro
UPSTREAM: wireguard: queueing: do not account for pfmemalloc when clearing skb header
UPSTREAM: wireguard: selftests: remove ancient kernel compatibility code
UPSTREAM: wireguard: allowedips: use kfree_rcu() instead of call_rcu()
UPSTREAM: wireguard: main: remove unused include <linux/version.h>
ANDROID: GKI: enable CONFIG_WIREGUARD
UPSTREAM: wireguard: global: fix spelling mistakes in comments
UPSTREAM: wireguard: Kconfig: select parent dependency for crypto
UPSTREAM: wireguard: selftests: import harness makefile for test suite
lib/crypto: blake2s: move hmac construction into wireguard
UPSTREAM: net: introduce skb_list_walk_safe for skb segment walking
UPSTREAM: net: WireGuard secure network tunnel
UPSTREAM: crypto: poly1305-x86_64 - Use XORL r32,32
UPSTREAM: crypto: curve25519-x86_64 - Use XORL r32,32
UPSTREAM: crypto: arm/poly1305 - Add prototype for poly1305_blocks_neon
UPSTREAM: crypto: arm/curve25519 - include <linux/scatterlist.h>
UPSTREAM: crypto: x86/curve25519 - Remove unused carry variables
UPSTREAM: crypto: x86/chacha-sse3 - use unaligned loads for state array
UPSTREAM: crypto: lib/chacha20poly1305 - Add missing function declaration
UPSTREAM: crypto: arch/lib - limit simd usage to 4k chunks
UPSTREAM: crypto: arm[64]/poly1305 - add artifact to .gitignore files
UPSTREAM: crypto: x86/curve25519 - leave r12 as spare register
UPSTREAM: crypto: x86/curve25519 - replace with formally verified implementation
UPSTREAM: crypto: arm64/chacha - correctly walk through blocks
UPSTREAM: crypto: x86/curve25519 - support assemblers with no adx support
UPSTREAM: crypto: chacha20poly1305 - prevent integer overflow on large input
UPSTREAM: crypto: Kconfig - allow tests to be disabled when manager is disabled
UPSTREAM: crypto: arm/chacha - fix build failured when kernel mode NEON is disabled
UPSTREAM: crypto: x86/poly1305 - emit does base conversion itself
UPSTREAM: crypto: chacha20poly1305 - add back missing test vectors and test chunking
UPSTREAM: crypto: x86/poly1305 - fix .gitignore typo
UPSTREAM: crypto: curve25519 - Fix selftest build error
UPSTREAM: crypto: {arm,arm64,mips}/poly1305 - remove redundant non-reduction from emit
UPSTREAM: crypto: x86/poly1305 - wire up faster implementations for kernel
UPSTREAM: crypto: x86/poly1305 - import unmodified cryptogams implementation
UPSTREAM: crypto: poly1305 - add new 32 and 64-bit generic versions
UPSTREAM: crypto: lib/curve25519 - re-add selftests
UPSTREAM: crypto: arm/curve25519 - add arch-specific key generation function
UPSTREAM: crypto: chacha - fix warning message in header file
BACKPORT: crypto: arch - conditionalize crypto api in arch glue for lib code
UPSTREAM: crypto: lib/chacha20poly1305 - use chacha20_crypt()
UPSTREAM: crypto: x86/chacha - only unregister algorithms if registered
UPSTREAM: crypto: chacha_generic - remove unnecessary setkey() functions
UPSTREAM: crypto: lib/chacha20poly1305 - reimplement crypt_from_sg() routine
UPSTREAM: crypto: chacha20poly1305 - import construction and selftest from Zinc
UPSTREAM: crypto: arm/curve25519 - wire up NEON implementation
UPSTREAM: crypto: arm/curve25519 - import Bernstein and Schwabe's Curve25519 ARM implementation
UPSTREAM: crypto: curve25519 - x86_64 library and KPP implementations
UPSTREAM: crypto: lib/curve25519 - work around Clang stack spilling issue
UPSTREAM: crypto: curve25519 - implement generic KPP driver
UPSTREAM: crypto: curve25519 - add kpp selftest
UPSTREAM: crypto: curve25519 - generic C library implementations
UPSTREAM: crypto: mips/poly1305 - incorporate OpenSSL/CRYPTOGAMS optimized implementation
UPSTREAM: crypto: arm/poly1305 - incorporate OpenSSL/CRYPTOGAMS NEON implementation
UPSTREAM: crypto: arm64/poly1305 - incorporate OpenSSL/CRYPTOGAMS NEON implementation
UPSTREAM: crypto: x86/poly1305 - expose existing driver as poly1305 library
UPSTREAM: crypto: x86/poly1305 - depend on generic library not generic shash
UPSTREAM: crypto: poly1305 - expose init/update/final library interface
UPSTREAM: crypto: x86/poly1305 - unify Poly1305 state struct with generic code
UPSTREAM: crypto: poly1305 - move core routines into a separate library
UPSTREAM: crypto: chacha - unexport chacha_generic routines
UPSTREAM: crypto: mips/chacha - wire up accelerated 32r2 code from Zinc
UPSTREAM: crypto: mips/chacha - import 32r2 ChaCha code from Zinc
UPSTREAM: crypto: arm/chacha - expose ARM ChaCha routine as library function
UPSTREAM: crypto: arm/chacha - remove dependency on generic ChaCha driver
UPSTREAM: crypto: arm/chacha - import Eric Biggers's scalar accelerated ChaCha code
UPSTREAM: crypto: arm64/chacha - expose arm64 ChaCha routine as library function
UPSTREAM: crypto: arm64/chacha - depend on generic chacha library instead of crypto driver
UPSTREAM: crypto: x86/chacha - expose SIMD ChaCha routine as library function
UPSTREAM: crypto: x86/chacha - depend on generic chacha library instead of crypto driver
UPSTREAM: crypto: chacha - move existing library code into lib/crypto
Revert "BACKPORT: crypto: arch - conditionalize crypto api in arch glue for lib code"
-----
Change-Id: I15633145724b6409932a3a440ca8e55e186d12c1
Signed-off-by: Alexander Martinz <amartinz@shiftphones.com>
This commit is contained in:
commit
c81e05c469
135 changed files with 40304 additions and 1938 deletions
|
|
@ -17595,6 +17595,14 @@ L: linux-gpio@vger.kernel.org
|
|||
S: Maintained
|
||||
F: drivers/gpio/gpio-ws16c48.c
|
||||
|
||||
WIREGUARD SECURE NETWORK TUNNEL
|
||||
M: Jason A. Donenfeld <Jason@zx2c4.com>
|
||||
S: Maintained
|
||||
F: drivers/net/wireguard/
|
||||
F: tools/testing/selftests/wireguard/
|
||||
L: wireguard@lists.zx2c4.com
|
||||
L: netdev@vger.kernel.org
|
||||
|
||||
WISTRON LAPTOP BUTTON DRIVER
|
||||
M: Miloslav Trmac <mitr@volny.cz>
|
||||
S: Maintained
|
||||
|
|
|
|||
1
arch/arm/crypto/.gitignore
vendored
1
arch/arm/crypto/.gitignore
vendored
|
|
@ -1,3 +1,4 @@
|
|||
aesbs-core.S
|
||||
sha256-core.S
|
||||
sha512-core.S
|
||||
poly1305-core.S
|
||||
|
|
|
|||
|
|
@ -148,14 +148,24 @@ config CRYPTO_CRC32_ARM_CE
|
|||
select CRYPTO_HASH
|
||||
|
||||
config CRYPTO_CHACHA20_NEON
|
||||
tristate "NEON accelerated ChaCha stream cipher algorithms"
|
||||
depends on KERNEL_MODE_NEON
|
||||
tristate "NEON and scalar accelerated ChaCha stream cipher algorithms"
|
||||
select CRYPTO_BLKCIPHER
|
||||
select CRYPTO_CHACHA20
|
||||
select CRYPTO_ARCH_HAVE_LIB_CHACHA
|
||||
|
||||
config CRYPTO_POLY1305_ARM
|
||||
tristate "Accelerated scalar and SIMD Poly1305 hash implementations"
|
||||
select CRYPTO_HASH
|
||||
select CRYPTO_ARCH_HAVE_LIB_POLY1305
|
||||
|
||||
config CRYPTO_NHPOLY1305_NEON
|
||||
tristate "NEON accelerated NHPoly1305 hash function (for Adiantum)"
|
||||
depends on KERNEL_MODE_NEON
|
||||
select CRYPTO_NHPOLY1305
|
||||
|
||||
config CRYPTO_CURVE25519_NEON
|
||||
tristate "NEON accelerated Curve25519 scalar multiplication library"
|
||||
depends on KERNEL_MODE_NEON
|
||||
select CRYPTO_LIB_CURVE25519_GENERIC
|
||||
select CRYPTO_ARCH_HAVE_LIB_CURVE25519
|
||||
|
||||
endif
|
||||
|
|
|
|||
|
|
@ -13,7 +13,9 @@ obj-$(CONFIG_CRYPTO_BLAKE2S_ARM) += blake2s-arm.o
|
|||
obj-$(if $(CONFIG_CRYPTO_BLAKE2S_ARM),y) += libblake2s-arm.o
|
||||
obj-$(CONFIG_CRYPTO_BLAKE2B_NEON) += blake2b-neon.o
|
||||
obj-$(CONFIG_CRYPTO_CHACHA20_NEON) += chacha-neon.o
|
||||
obj-$(CONFIG_CRYPTO_POLY1305_ARM) += poly1305-arm.o
|
||||
obj-$(CONFIG_CRYPTO_NHPOLY1305_NEON) += nhpoly1305-neon.o
|
||||
obj-$(CONFIG_CRYPTO_CURVE25519_NEON) += curve25519-neon.o
|
||||
|
||||
obj-$(CONFIG_CRYPTO_AES_ARM_CE) += aes-arm-ce.o
|
||||
obj-$(CONFIG_CRYPTO_SHA1_ARM_CE) += sha1-arm-ce.o
|
||||
|
|
@ -39,13 +41,19 @@ aes-arm-ce-y := aes-ce-core.o aes-ce-glue.o
|
|||
ghash-arm-ce-y := ghash-ce-core.o ghash-ce-glue.o
|
||||
crct10dif-arm-ce-y := crct10dif-ce-core.o crct10dif-ce-glue.o
|
||||
crc32-arm-ce-y:= crc32-ce-core.o crc32-ce-glue.o
|
||||
chacha-neon-y := chacha-neon-core.o chacha-neon-glue.o
|
||||
chacha-neon-y := chacha-scalar-core.o chacha-glue.o
|
||||
chacha-neon-$(CONFIG_KERNEL_MODE_NEON) += chacha-neon-core.o
|
||||
poly1305-arm-y := poly1305-core.o poly1305-glue.o
|
||||
nhpoly1305-neon-y := nh-neon-core.o nhpoly1305-neon-glue.o
|
||||
curve25519-neon-y := curve25519-core.o curve25519-glue.o
|
||||
|
||||
ifdef REGENERATE_ARM_CRYPTO
|
||||
quiet_cmd_perl = PERL $@
|
||||
cmd_perl = $(PERL) $(<) > $(@)
|
||||
|
||||
$(src)/poly1305-core.S_shipped: $(src)/poly1305-armv4.pl
|
||||
$(call cmd,perl)
|
||||
|
||||
$(src)/sha256-core.S_shipped: $(src)/sha256-armv4.pl
|
||||
$(call cmd,perl)
|
||||
|
||||
|
|
@ -53,4 +61,9 @@ $(src)/sha512-core.S_shipped: $(src)/sha512-armv4.pl
|
|||
$(call cmd,perl)
|
||||
endif
|
||||
|
||||
clean-files += sha256-core.S sha512-core.S
|
||||
clean-files += poly1305-core.S sha256-core.S sha512-core.S
|
||||
|
||||
# massage the perlasm code a bit so we only get the NEON routine if we need it
|
||||
poly1305-aflags-$(CONFIG_CPU_V7) := -U__LINUX_ARM_ARCH__ -D__LINUX_ARM_ARCH__=5
|
||||
poly1305-aflags-$(CONFIG_KERNEL_MODE_NEON) := -U__LINUX_ARM_ARCH__ -D__LINUX_ARM_ARCH__=7
|
||||
AFLAGS_poly1305-core.o += $(poly1305-aflags-y)
|
||||
|
|
|
|||
357
arch/arm/crypto/chacha-glue.c
Normal file
357
arch/arm/crypto/chacha-glue.c
Normal file
|
|
@ -0,0 +1,357 @@
|
|||
// SPDX-License-Identifier: GPL-2.0
|
||||
/*
|
||||
* ARM NEON accelerated ChaCha and XChaCha stream ciphers,
|
||||
* including ChaCha20 (RFC7539)
|
||||
*
|
||||
* Copyright (C) 2016-2019 Linaro, Ltd. <ard.biesheuvel@linaro.org>
|
||||
* Copyright (C) 2015 Martin Willi
|
||||
*/
|
||||
|
||||
#include <crypto/algapi.h>
|
||||
#include <crypto/internal/chacha.h>
|
||||
#include <crypto/internal/simd.h>
|
||||
#include <crypto/internal/skcipher.h>
|
||||
#include <linux/jump_label.h>
|
||||
#include <linux/kernel.h>
|
||||
#include <linux/module.h>
|
||||
|
||||
#include <asm/cputype.h>
|
||||
#include <asm/hwcap.h>
|
||||
#include <asm/neon.h>
|
||||
#include <asm/simd.h>
|
||||
|
||||
asmlinkage void chacha_block_xor_neon(const u32 *state, u8 *dst, const u8 *src,
|
||||
int nrounds);
|
||||
asmlinkage void chacha_4block_xor_neon(const u32 *state, u8 *dst, const u8 *src,
|
||||
int nrounds);
|
||||
asmlinkage void hchacha_block_arm(const u32 *state, u32 *out, int nrounds);
|
||||
asmlinkage void hchacha_block_neon(const u32 *state, u32 *out, int nrounds);
|
||||
|
||||
asmlinkage void chacha_doarm(u8 *dst, const u8 *src, unsigned int bytes,
|
||||
const u32 *state, int nrounds);
|
||||
|
||||
static __ro_after_init DEFINE_STATIC_KEY_FALSE(use_neon);
|
||||
|
||||
static inline bool neon_usable(void)
|
||||
{
|
||||
return static_branch_likely(&use_neon) && crypto_simd_usable();
|
||||
}
|
||||
|
||||
static void chacha_doneon(u32 *state, u8 *dst, const u8 *src,
|
||||
unsigned int bytes, int nrounds)
|
||||
{
|
||||
u8 buf[CHACHA_BLOCK_SIZE];
|
||||
|
||||
while (bytes >= CHACHA_BLOCK_SIZE * 4) {
|
||||
chacha_4block_xor_neon(state, dst, src, nrounds);
|
||||
bytes -= CHACHA_BLOCK_SIZE * 4;
|
||||
src += CHACHA_BLOCK_SIZE * 4;
|
||||
dst += CHACHA_BLOCK_SIZE * 4;
|
||||
state[12] += 4;
|
||||
}
|
||||
while (bytes >= CHACHA_BLOCK_SIZE) {
|
||||
chacha_block_xor_neon(state, dst, src, nrounds);
|
||||
bytes -= CHACHA_BLOCK_SIZE;
|
||||
src += CHACHA_BLOCK_SIZE;
|
||||
dst += CHACHA_BLOCK_SIZE;
|
||||
state[12]++;
|
||||
}
|
||||
if (bytes) {
|
||||
memcpy(buf, src, bytes);
|
||||
chacha_block_xor_neon(state, buf, buf, nrounds);
|
||||
memcpy(dst, buf, bytes);
|
||||
}
|
||||
}
|
||||
|
||||
void hchacha_block_arch(const u32 *state, u32 *stream, int nrounds)
|
||||
{
|
||||
if (!IS_ENABLED(CONFIG_KERNEL_MODE_NEON) || !neon_usable()) {
|
||||
hchacha_block_arm(state, stream, nrounds);
|
||||
} else {
|
||||
kernel_neon_begin();
|
||||
hchacha_block_neon(state, stream, nrounds);
|
||||
kernel_neon_end();
|
||||
}
|
||||
}
|
||||
EXPORT_SYMBOL(hchacha_block_arch);
|
||||
|
||||
void chacha_init_arch(u32 *state, const u32 *key, const u8 *iv)
|
||||
{
|
||||
chacha_init_generic(state, key, iv);
|
||||
}
|
||||
EXPORT_SYMBOL(chacha_init_arch);
|
||||
|
||||
void chacha_crypt_arch(u32 *state, u8 *dst, const u8 *src, unsigned int bytes,
|
||||
int nrounds)
|
||||
{
|
||||
if (!IS_ENABLED(CONFIG_KERNEL_MODE_NEON) || !neon_usable() ||
|
||||
bytes <= CHACHA_BLOCK_SIZE) {
|
||||
chacha_doarm(dst, src, bytes, state, nrounds);
|
||||
state[12] += DIV_ROUND_UP(bytes, CHACHA_BLOCK_SIZE);
|
||||
return;
|
||||
}
|
||||
|
||||
do {
|
||||
unsigned int todo = min_t(unsigned int, bytes, SZ_4K);
|
||||
|
||||
kernel_neon_begin();
|
||||
chacha_doneon(state, dst, src, todo, nrounds);
|
||||
kernel_neon_end();
|
||||
|
||||
bytes -= todo;
|
||||
src += todo;
|
||||
dst += todo;
|
||||
} while (bytes);
|
||||
}
|
||||
EXPORT_SYMBOL(chacha_crypt_arch);
|
||||
|
||||
static int chacha_stream_xor(struct skcipher_request *req,
|
||||
const struct chacha_ctx *ctx, const u8 *iv,
|
||||
bool neon)
|
||||
{
|
||||
struct skcipher_walk walk;
|
||||
u32 state[16];
|
||||
int err;
|
||||
|
||||
err = skcipher_walk_virt(&walk, req, false);
|
||||
|
||||
chacha_init_generic(state, ctx->key, iv);
|
||||
|
||||
while (walk.nbytes > 0) {
|
||||
unsigned int nbytes = walk.nbytes;
|
||||
|
||||
if (nbytes < walk.total)
|
||||
nbytes = round_down(nbytes, walk.stride);
|
||||
|
||||
if (!IS_ENABLED(CONFIG_KERNEL_MODE_NEON) || !neon) {
|
||||
chacha_doarm(walk.dst.virt.addr, walk.src.virt.addr,
|
||||
nbytes, state, ctx->nrounds);
|
||||
state[12] += DIV_ROUND_UP(nbytes, CHACHA_BLOCK_SIZE);
|
||||
} else {
|
||||
kernel_neon_begin();
|
||||
chacha_doneon(state, walk.dst.virt.addr,
|
||||
walk.src.virt.addr, nbytes, ctx->nrounds);
|
||||
kernel_neon_end();
|
||||
}
|
||||
err = skcipher_walk_done(&walk, walk.nbytes - nbytes);
|
||||
}
|
||||
|
||||
return err;
|
||||
}
|
||||
|
||||
static int do_chacha(struct skcipher_request *req, bool neon)
|
||||
{
|
||||
struct crypto_skcipher *tfm = crypto_skcipher_reqtfm(req);
|
||||
struct chacha_ctx *ctx = crypto_skcipher_ctx(tfm);
|
||||
|
||||
return chacha_stream_xor(req, ctx, req->iv, neon);
|
||||
}
|
||||
|
||||
static int chacha_arm(struct skcipher_request *req)
|
||||
{
|
||||
return do_chacha(req, false);
|
||||
}
|
||||
|
||||
static int chacha_neon(struct skcipher_request *req)
|
||||
{
|
||||
return do_chacha(req, neon_usable());
|
||||
}
|
||||
|
||||
static int do_xchacha(struct skcipher_request *req, bool neon)
|
||||
{
|
||||
struct crypto_skcipher *tfm = crypto_skcipher_reqtfm(req);
|
||||
struct chacha_ctx *ctx = crypto_skcipher_ctx(tfm);
|
||||
struct chacha_ctx subctx;
|
||||
u32 state[16];
|
||||
u8 real_iv[16];
|
||||
|
||||
chacha_init_generic(state, ctx->key, req->iv);
|
||||
|
||||
if (!IS_ENABLED(CONFIG_KERNEL_MODE_NEON) || !neon) {
|
||||
hchacha_block_arm(state, subctx.key, ctx->nrounds);
|
||||
} else {
|
||||
kernel_neon_begin();
|
||||
hchacha_block_neon(state, subctx.key, ctx->nrounds);
|
||||
kernel_neon_end();
|
||||
}
|
||||
subctx.nrounds = ctx->nrounds;
|
||||
|
||||
memcpy(&real_iv[0], req->iv + 24, 8);
|
||||
memcpy(&real_iv[8], req->iv + 16, 8);
|
||||
return chacha_stream_xor(req, &subctx, real_iv, neon);
|
||||
}
|
||||
|
||||
static int xchacha_arm(struct skcipher_request *req)
|
||||
{
|
||||
return do_xchacha(req, false);
|
||||
}
|
||||
|
||||
static int xchacha_neon(struct skcipher_request *req)
|
||||
{
|
||||
return do_xchacha(req, neon_usable());
|
||||
}
|
||||
|
||||
static struct skcipher_alg arm_algs[] = {
|
||||
{
|
||||
.base.cra_name = "chacha20",
|
||||
.base.cra_driver_name = "chacha20-arm",
|
||||
.base.cra_priority = 200,
|
||||
.base.cra_blocksize = 1,
|
||||
.base.cra_ctxsize = sizeof(struct chacha_ctx),
|
||||
.base.cra_module = THIS_MODULE,
|
||||
|
||||
.min_keysize = CHACHA_KEY_SIZE,
|
||||
.max_keysize = CHACHA_KEY_SIZE,
|
||||
.ivsize = CHACHA_IV_SIZE,
|
||||
.chunksize = CHACHA_BLOCK_SIZE,
|
||||
.setkey = chacha20_setkey,
|
||||
.encrypt = chacha_arm,
|
||||
.decrypt = chacha_arm,
|
||||
}, {
|
||||
.base.cra_name = "xchacha20",
|
||||
.base.cra_driver_name = "xchacha20-arm",
|
||||
.base.cra_priority = 200,
|
||||
.base.cra_blocksize = 1,
|
||||
.base.cra_ctxsize = sizeof(struct chacha_ctx),
|
||||
.base.cra_module = THIS_MODULE,
|
||||
|
||||
.min_keysize = CHACHA_KEY_SIZE,
|
||||
.max_keysize = CHACHA_KEY_SIZE,
|
||||
.ivsize = XCHACHA_IV_SIZE,
|
||||
.chunksize = CHACHA_BLOCK_SIZE,
|
||||
.setkey = chacha20_setkey,
|
||||
.encrypt = xchacha_arm,
|
||||
.decrypt = xchacha_arm,
|
||||
}, {
|
||||
.base.cra_name = "xchacha12",
|
||||
.base.cra_driver_name = "xchacha12-arm",
|
||||
.base.cra_priority = 200,
|
||||
.base.cra_blocksize = 1,
|
||||
.base.cra_ctxsize = sizeof(struct chacha_ctx),
|
||||
.base.cra_module = THIS_MODULE,
|
||||
|
||||
.min_keysize = CHACHA_KEY_SIZE,
|
||||
.max_keysize = CHACHA_KEY_SIZE,
|
||||
.ivsize = XCHACHA_IV_SIZE,
|
||||
.chunksize = CHACHA_BLOCK_SIZE,
|
||||
.setkey = chacha12_setkey,
|
||||
.encrypt = xchacha_arm,
|
||||
.decrypt = xchacha_arm,
|
||||
},
|
||||
};
|
||||
|
||||
static struct skcipher_alg neon_algs[] = {
|
||||
{
|
||||
.base.cra_name = "chacha20",
|
||||
.base.cra_driver_name = "chacha20-neon",
|
||||
.base.cra_priority = 300,
|
||||
.base.cra_blocksize = 1,
|
||||
.base.cra_ctxsize = sizeof(struct chacha_ctx),
|
||||
.base.cra_module = THIS_MODULE,
|
||||
|
||||
.min_keysize = CHACHA_KEY_SIZE,
|
||||
.max_keysize = CHACHA_KEY_SIZE,
|
||||
.ivsize = CHACHA_IV_SIZE,
|
||||
.chunksize = CHACHA_BLOCK_SIZE,
|
||||
.walksize = 4 * CHACHA_BLOCK_SIZE,
|
||||
.setkey = chacha20_setkey,
|
||||
.encrypt = chacha_neon,
|
||||
.decrypt = chacha_neon,
|
||||
}, {
|
||||
.base.cra_name = "xchacha20",
|
||||
.base.cra_driver_name = "xchacha20-neon",
|
||||
.base.cra_priority = 300,
|
||||
.base.cra_blocksize = 1,
|
||||
.base.cra_ctxsize = sizeof(struct chacha_ctx),
|
||||
.base.cra_module = THIS_MODULE,
|
||||
|
||||
.min_keysize = CHACHA_KEY_SIZE,
|
||||
.max_keysize = CHACHA_KEY_SIZE,
|
||||
.ivsize = XCHACHA_IV_SIZE,
|
||||
.chunksize = CHACHA_BLOCK_SIZE,
|
||||
.walksize = 4 * CHACHA_BLOCK_SIZE,
|
||||
.setkey = chacha20_setkey,
|
||||
.encrypt = xchacha_neon,
|
||||
.decrypt = xchacha_neon,
|
||||
}, {
|
||||
.base.cra_name = "xchacha12",
|
||||
.base.cra_driver_name = "xchacha12-neon",
|
||||
.base.cra_priority = 300,
|
||||
.base.cra_blocksize = 1,
|
||||
.base.cra_ctxsize = sizeof(struct chacha_ctx),
|
||||
.base.cra_module = THIS_MODULE,
|
||||
|
||||
.min_keysize = CHACHA_KEY_SIZE,
|
||||
.max_keysize = CHACHA_KEY_SIZE,
|
||||
.ivsize = XCHACHA_IV_SIZE,
|
||||
.chunksize = CHACHA_BLOCK_SIZE,
|
||||
.walksize = 4 * CHACHA_BLOCK_SIZE,
|
||||
.setkey = chacha12_setkey,
|
||||
.encrypt = xchacha_neon,
|
||||
.decrypt = xchacha_neon,
|
||||
}
|
||||
};
|
||||
|
||||
static int __init chacha_simd_mod_init(void)
|
||||
{
|
||||
int err = 0;
|
||||
|
||||
if (IS_REACHABLE(CONFIG_CRYPTO_BLKCIPHER)) {
|
||||
err = crypto_register_skciphers(arm_algs, ARRAY_SIZE(arm_algs));
|
||||
if (err)
|
||||
return err;
|
||||
}
|
||||
|
||||
if (IS_ENABLED(CONFIG_KERNEL_MODE_NEON) && (elf_hwcap & HWCAP_NEON)) {
|
||||
int i;
|
||||
|
||||
switch (read_cpuid_part()) {
|
||||
case ARM_CPU_PART_CORTEX_A7:
|
||||
case ARM_CPU_PART_CORTEX_A5:
|
||||
/*
|
||||
* The Cortex-A7 and Cortex-A5 do not perform well with
|
||||
* the NEON implementation but do incredibly with the
|
||||
* scalar one and use less power.
|
||||
*/
|
||||
for (i = 0; i < ARRAY_SIZE(neon_algs); i++)
|
||||
neon_algs[i].base.cra_priority = 0;
|
||||
break;
|
||||
default:
|
||||
static_branch_enable(&use_neon);
|
||||
}
|
||||
|
||||
if (IS_REACHABLE(CONFIG_CRYPTO_BLKCIPHER)) {
|
||||
err = crypto_register_skciphers(neon_algs, ARRAY_SIZE(neon_algs));
|
||||
if (err)
|
||||
crypto_unregister_skciphers(arm_algs, ARRAY_SIZE(arm_algs));
|
||||
}
|
||||
}
|
||||
return err;
|
||||
}
|
||||
|
||||
static void __exit chacha_simd_mod_fini(void)
|
||||
{
|
||||
if (IS_REACHABLE(CONFIG_CRYPTO_BLKCIPHER)) {
|
||||
crypto_unregister_skciphers(arm_algs, ARRAY_SIZE(arm_algs));
|
||||
if (IS_ENABLED(CONFIG_KERNEL_MODE_NEON) && (elf_hwcap & HWCAP_NEON))
|
||||
crypto_unregister_skciphers(neon_algs, ARRAY_SIZE(neon_algs));
|
||||
}
|
||||
}
|
||||
|
||||
module_init(chacha_simd_mod_init);
|
||||
module_exit(chacha_simd_mod_fini);
|
||||
|
||||
MODULE_DESCRIPTION("ChaCha and XChaCha stream ciphers (scalar and NEON accelerated)");
|
||||
MODULE_AUTHOR("Ard Biesheuvel <ard.biesheuvel@linaro.org>");
|
||||
MODULE_LICENSE("GPL v2");
|
||||
MODULE_ALIAS_CRYPTO("chacha20");
|
||||
MODULE_ALIAS_CRYPTO("chacha20-arm");
|
||||
MODULE_ALIAS_CRYPTO("xchacha20");
|
||||
MODULE_ALIAS_CRYPTO("xchacha20-arm");
|
||||
MODULE_ALIAS_CRYPTO("xchacha12");
|
||||
MODULE_ALIAS_CRYPTO("xchacha12-arm");
|
||||
#ifdef CONFIG_KERNEL_MODE_NEON
|
||||
MODULE_ALIAS_CRYPTO("chacha20-neon");
|
||||
MODULE_ALIAS_CRYPTO("xchacha20-neon");
|
||||
MODULE_ALIAS_CRYPTO("xchacha12-neon");
|
||||
#endif
|
||||
|
|
@ -1,202 +0,0 @@
|
|||
/*
|
||||
* ARM NEON accelerated ChaCha and XChaCha stream ciphers,
|
||||
* including ChaCha20 (RFC7539)
|
||||
*
|
||||
* Copyright (C) 2016 Linaro, Ltd. <ard.biesheuvel@linaro.org>
|
||||
*
|
||||
* This program is free software; you can redistribute it and/or modify
|
||||
* it under the terms of the GNU General Public License version 2 as
|
||||
* published by the Free Software Foundation.
|
||||
*
|
||||
* Based on:
|
||||
* ChaCha20 256-bit cipher algorithm, RFC7539, SIMD glue code
|
||||
*
|
||||
* Copyright (C) 2015 Martin Willi
|
||||
*
|
||||
* This program is free software; you can redistribute it and/or modify
|
||||
* it under the terms of the GNU General Public License as published by
|
||||
* the Free Software Foundation; either version 2 of the License, or
|
||||
* (at your option) any later version.
|
||||
*/
|
||||
|
||||
#include <crypto/algapi.h>
|
||||
#include <crypto/chacha.h>
|
||||
#include <crypto/internal/simd.h>
|
||||
#include <crypto/internal/skcipher.h>
|
||||
#include <linux/kernel.h>
|
||||
#include <linux/module.h>
|
||||
|
||||
#include <asm/hwcap.h>
|
||||
#include <asm/neon.h>
|
||||
#include <asm/simd.h>
|
||||
|
||||
asmlinkage void chacha_block_xor_neon(const u32 *state, u8 *dst, const u8 *src,
|
||||
int nrounds);
|
||||
asmlinkage void chacha_4block_xor_neon(const u32 *state, u8 *dst, const u8 *src,
|
||||
int nrounds);
|
||||
asmlinkage void hchacha_block_neon(const u32 *state, u32 *out, int nrounds);
|
||||
|
||||
static void chacha_doneon(u32 *state, u8 *dst, const u8 *src,
|
||||
unsigned int bytes, int nrounds)
|
||||
{
|
||||
u8 buf[CHACHA_BLOCK_SIZE];
|
||||
|
||||
while (bytes >= CHACHA_BLOCK_SIZE * 4) {
|
||||
chacha_4block_xor_neon(state, dst, src, nrounds);
|
||||
bytes -= CHACHA_BLOCK_SIZE * 4;
|
||||
src += CHACHA_BLOCK_SIZE * 4;
|
||||
dst += CHACHA_BLOCK_SIZE * 4;
|
||||
state[12] += 4;
|
||||
}
|
||||
while (bytes >= CHACHA_BLOCK_SIZE) {
|
||||
chacha_block_xor_neon(state, dst, src, nrounds);
|
||||
bytes -= CHACHA_BLOCK_SIZE;
|
||||
src += CHACHA_BLOCK_SIZE;
|
||||
dst += CHACHA_BLOCK_SIZE;
|
||||
state[12]++;
|
||||
}
|
||||
if (bytes) {
|
||||
memcpy(buf, src, bytes);
|
||||
chacha_block_xor_neon(state, buf, buf, nrounds);
|
||||
memcpy(dst, buf, bytes);
|
||||
}
|
||||
}
|
||||
|
||||
static int chacha_neon_stream_xor(struct skcipher_request *req,
|
||||
const struct chacha_ctx *ctx, const u8 *iv)
|
||||
{
|
||||
struct skcipher_walk walk;
|
||||
u32 state[16];
|
||||
int err;
|
||||
|
||||
err = skcipher_walk_virt(&walk, req, false);
|
||||
|
||||
crypto_chacha_init(state, ctx, iv);
|
||||
|
||||
while (walk.nbytes > 0) {
|
||||
unsigned int nbytes = walk.nbytes;
|
||||
|
||||
if (nbytes < walk.total)
|
||||
nbytes = round_down(nbytes, walk.stride);
|
||||
|
||||
kernel_neon_begin();
|
||||
chacha_doneon(state, walk.dst.virt.addr, walk.src.virt.addr,
|
||||
nbytes, ctx->nrounds);
|
||||
kernel_neon_end();
|
||||
err = skcipher_walk_done(&walk, walk.nbytes - nbytes);
|
||||
}
|
||||
|
||||
return err;
|
||||
}
|
||||
|
||||
static int chacha_neon(struct skcipher_request *req)
|
||||
{
|
||||
struct crypto_skcipher *tfm = crypto_skcipher_reqtfm(req);
|
||||
struct chacha_ctx *ctx = crypto_skcipher_ctx(tfm);
|
||||
|
||||
if (req->cryptlen <= CHACHA_BLOCK_SIZE || !crypto_simd_usable())
|
||||
return crypto_chacha_crypt(req);
|
||||
|
||||
return chacha_neon_stream_xor(req, ctx, req->iv);
|
||||
}
|
||||
|
||||
static int xchacha_neon(struct skcipher_request *req)
|
||||
{
|
||||
struct crypto_skcipher *tfm = crypto_skcipher_reqtfm(req);
|
||||
struct chacha_ctx *ctx = crypto_skcipher_ctx(tfm);
|
||||
struct chacha_ctx subctx;
|
||||
u32 state[16];
|
||||
u8 real_iv[16];
|
||||
|
||||
if (req->cryptlen <= CHACHA_BLOCK_SIZE || !crypto_simd_usable())
|
||||
return crypto_xchacha_crypt(req);
|
||||
|
||||
crypto_chacha_init(state, ctx, req->iv);
|
||||
|
||||
kernel_neon_begin();
|
||||
hchacha_block_neon(state, subctx.key, ctx->nrounds);
|
||||
kernel_neon_end();
|
||||
subctx.nrounds = ctx->nrounds;
|
||||
|
||||
memcpy(&real_iv[0], req->iv + 24, 8);
|
||||
memcpy(&real_iv[8], req->iv + 16, 8);
|
||||
return chacha_neon_stream_xor(req, &subctx, real_iv);
|
||||
}
|
||||
|
||||
static struct skcipher_alg algs[] = {
|
||||
{
|
||||
.base.cra_name = "chacha20",
|
||||
.base.cra_driver_name = "chacha20-neon",
|
||||
.base.cra_priority = 300,
|
||||
.base.cra_blocksize = 1,
|
||||
.base.cra_ctxsize = sizeof(struct chacha_ctx),
|
||||
.base.cra_module = THIS_MODULE,
|
||||
|
||||
.min_keysize = CHACHA_KEY_SIZE,
|
||||
.max_keysize = CHACHA_KEY_SIZE,
|
||||
.ivsize = CHACHA_IV_SIZE,
|
||||
.chunksize = CHACHA_BLOCK_SIZE,
|
||||
.walksize = 4 * CHACHA_BLOCK_SIZE,
|
||||
.setkey = crypto_chacha20_setkey,
|
||||
.encrypt = chacha_neon,
|
||||
.decrypt = chacha_neon,
|
||||
}, {
|
||||
.base.cra_name = "xchacha20",
|
||||
.base.cra_driver_name = "xchacha20-neon",
|
||||
.base.cra_priority = 300,
|
||||
.base.cra_blocksize = 1,
|
||||
.base.cra_ctxsize = sizeof(struct chacha_ctx),
|
||||
.base.cra_module = THIS_MODULE,
|
||||
|
||||
.min_keysize = CHACHA_KEY_SIZE,
|
||||
.max_keysize = CHACHA_KEY_SIZE,
|
||||
.ivsize = XCHACHA_IV_SIZE,
|
||||
.chunksize = CHACHA_BLOCK_SIZE,
|
||||
.walksize = 4 * CHACHA_BLOCK_SIZE,
|
||||
.setkey = crypto_chacha20_setkey,
|
||||
.encrypt = xchacha_neon,
|
||||
.decrypt = xchacha_neon,
|
||||
}, {
|
||||
.base.cra_name = "xchacha12",
|
||||
.base.cra_driver_name = "xchacha12-neon",
|
||||
.base.cra_priority = 300,
|
||||
.base.cra_blocksize = 1,
|
||||
.base.cra_ctxsize = sizeof(struct chacha_ctx),
|
||||
.base.cra_module = THIS_MODULE,
|
||||
|
||||
.min_keysize = CHACHA_KEY_SIZE,
|
||||
.max_keysize = CHACHA_KEY_SIZE,
|
||||
.ivsize = XCHACHA_IV_SIZE,
|
||||
.chunksize = CHACHA_BLOCK_SIZE,
|
||||
.walksize = 4 * CHACHA_BLOCK_SIZE,
|
||||
.setkey = crypto_chacha12_setkey,
|
||||
.encrypt = xchacha_neon,
|
||||
.decrypt = xchacha_neon,
|
||||
}
|
||||
};
|
||||
|
||||
static int __init chacha_simd_mod_init(void)
|
||||
{
|
||||
if (!(elf_hwcap & HWCAP_NEON))
|
||||
return -ENODEV;
|
||||
|
||||
return crypto_register_skciphers(algs, ARRAY_SIZE(algs));
|
||||
}
|
||||
|
||||
static void __exit chacha_simd_mod_fini(void)
|
||||
{
|
||||
crypto_unregister_skciphers(algs, ARRAY_SIZE(algs));
|
||||
}
|
||||
|
||||
module_init(chacha_simd_mod_init);
|
||||
module_exit(chacha_simd_mod_fini);
|
||||
|
||||
MODULE_DESCRIPTION("ChaCha and XChaCha stream ciphers (NEON accelerated)");
|
||||
MODULE_AUTHOR("Ard Biesheuvel <ard.biesheuvel@linaro.org>");
|
||||
MODULE_LICENSE("GPL v2");
|
||||
MODULE_ALIAS_CRYPTO("chacha20");
|
||||
MODULE_ALIAS_CRYPTO("chacha20-neon");
|
||||
MODULE_ALIAS_CRYPTO("xchacha20");
|
||||
MODULE_ALIAS_CRYPTO("xchacha20-neon");
|
||||
MODULE_ALIAS_CRYPTO("xchacha12");
|
||||
MODULE_ALIAS_CRYPTO("xchacha12-neon");
|
||||
460
arch/arm/crypto/chacha-scalar-core.S
Normal file
460
arch/arm/crypto/chacha-scalar-core.S
Normal file
|
|
@ -0,0 +1,460 @@
|
|||
/* SPDX-License-Identifier: GPL-2.0 */
|
||||
/*
|
||||
* Copyright (C) 2018 Google, Inc.
|
||||
*/
|
||||
|
||||
#include <linux/linkage.h>
|
||||
#include <asm/assembler.h>
|
||||
|
||||
/*
|
||||
* Design notes:
|
||||
*
|
||||
* 16 registers would be needed to hold the state matrix, but only 14 are
|
||||
* available because 'sp' and 'pc' cannot be used. So we spill the elements
|
||||
* (x8, x9) to the stack and swap them out with (x10, x11). This adds one
|
||||
* 'ldrd' and one 'strd' instruction per round.
|
||||
*
|
||||
* All rotates are performed using the implicit rotate operand accepted by the
|
||||
* 'add' and 'eor' instructions. This is faster than using explicit rotate
|
||||
* instructions. To make this work, we allow the values in the second and last
|
||||
* rows of the ChaCha state matrix (rows 'b' and 'd') to temporarily have the
|
||||
* wrong rotation amount. The rotation amount is then fixed up just in time
|
||||
* when the values are used. 'brot' is the number of bits the values in row 'b'
|
||||
* need to be rotated right to arrive at the correct values, and 'drot'
|
||||
* similarly for row 'd'. (brot, drot) start out as (0, 0) but we make it such
|
||||
* that they end up as (25, 24) after every round.
|
||||
*/
|
||||
|
||||
// ChaCha state registers
|
||||
X0 .req r0
|
||||
X1 .req r1
|
||||
X2 .req r2
|
||||
X3 .req r3
|
||||
X4 .req r4
|
||||
X5 .req r5
|
||||
X6 .req r6
|
||||
X7 .req r7
|
||||
X8_X10 .req r8 // shared by x8 and x10
|
||||
X9_X11 .req r9 // shared by x9 and x11
|
||||
X12 .req r10
|
||||
X13 .req r11
|
||||
X14 .req r12
|
||||
X15 .req r14
|
||||
|
||||
.macro __rev out, in, t0, t1, t2
|
||||
.if __LINUX_ARM_ARCH__ >= 6
|
||||
rev \out, \in
|
||||
.else
|
||||
lsl \t0, \in, #24
|
||||
and \t1, \in, #0xff00
|
||||
and \t2, \in, #0xff0000
|
||||
orr \out, \t0, \in, lsr #24
|
||||
orr \out, \out, \t1, lsl #8
|
||||
orr \out, \out, \t2, lsr #8
|
||||
.endif
|
||||
.endm
|
||||
|
||||
.macro _le32_bswap x, t0, t1, t2
|
||||
#ifdef __ARMEB__
|
||||
__rev \x, \x, \t0, \t1, \t2
|
||||
#endif
|
||||
.endm
|
||||
|
||||
.macro _le32_bswap_4x a, b, c, d, t0, t1, t2
|
||||
_le32_bswap \a, \t0, \t1, \t2
|
||||
_le32_bswap \b, \t0, \t1, \t2
|
||||
_le32_bswap \c, \t0, \t1, \t2
|
||||
_le32_bswap \d, \t0, \t1, \t2
|
||||
.endm
|
||||
|
||||
.macro __ldrd a, b, src, offset
|
||||
#if __LINUX_ARM_ARCH__ >= 6
|
||||
ldrd \a, \b, [\src, #\offset]
|
||||
#else
|
||||
ldr \a, [\src, #\offset]
|
||||
ldr \b, [\src, #\offset + 4]
|
||||
#endif
|
||||
.endm
|
||||
|
||||
.macro __strd a, b, dst, offset
|
||||
#if __LINUX_ARM_ARCH__ >= 6
|
||||
strd \a, \b, [\dst, #\offset]
|
||||
#else
|
||||
str \a, [\dst, #\offset]
|
||||
str \b, [\dst, #\offset + 4]
|
||||
#endif
|
||||
.endm
|
||||
|
||||
.macro _halfround a1, b1, c1, d1, a2, b2, c2, d2
|
||||
|
||||
// a += b; d ^= a; d = rol(d, 16);
|
||||
add \a1, \a1, \b1, ror #brot
|
||||
add \a2, \a2, \b2, ror #brot
|
||||
eor \d1, \a1, \d1, ror #drot
|
||||
eor \d2, \a2, \d2, ror #drot
|
||||
// drot == 32 - 16 == 16
|
||||
|
||||
// c += d; b ^= c; b = rol(b, 12);
|
||||
add \c1, \c1, \d1, ror #16
|
||||
add \c2, \c2, \d2, ror #16
|
||||
eor \b1, \c1, \b1, ror #brot
|
||||
eor \b2, \c2, \b2, ror #brot
|
||||
// brot == 32 - 12 == 20
|
||||
|
||||
// a += b; d ^= a; d = rol(d, 8);
|
||||
add \a1, \a1, \b1, ror #20
|
||||
add \a2, \a2, \b2, ror #20
|
||||
eor \d1, \a1, \d1, ror #16
|
||||
eor \d2, \a2, \d2, ror #16
|
||||
// drot == 32 - 8 == 24
|
||||
|
||||
// c += d; b ^= c; b = rol(b, 7);
|
||||
add \c1, \c1, \d1, ror #24
|
||||
add \c2, \c2, \d2, ror #24
|
||||
eor \b1, \c1, \b1, ror #20
|
||||
eor \b2, \c2, \b2, ror #20
|
||||
// brot == 32 - 7 == 25
|
||||
.endm
|
||||
|
||||
.macro _doubleround
|
||||
|
||||
// column round
|
||||
|
||||
// quarterrounds: (x0, x4, x8, x12) and (x1, x5, x9, x13)
|
||||
_halfround X0, X4, X8_X10, X12, X1, X5, X9_X11, X13
|
||||
|
||||
// save (x8, x9); restore (x10, x11)
|
||||
__strd X8_X10, X9_X11, sp, 0
|
||||
__ldrd X8_X10, X9_X11, sp, 8
|
||||
|
||||
// quarterrounds: (x2, x6, x10, x14) and (x3, x7, x11, x15)
|
||||
_halfround X2, X6, X8_X10, X14, X3, X7, X9_X11, X15
|
||||
|
||||
.set brot, 25
|
||||
.set drot, 24
|
||||
|
||||
// diagonal round
|
||||
|
||||
// quarterrounds: (x0, x5, x10, x15) and (x1, x6, x11, x12)
|
||||
_halfround X0, X5, X8_X10, X15, X1, X6, X9_X11, X12
|
||||
|
||||
// save (x10, x11); restore (x8, x9)
|
||||
__strd X8_X10, X9_X11, sp, 8
|
||||
__ldrd X8_X10, X9_X11, sp, 0
|
||||
|
||||
// quarterrounds: (x2, x7, x8, x13) and (x3, x4, x9, x14)
|
||||
_halfround X2, X7, X8_X10, X13, X3, X4, X9_X11, X14
|
||||
.endm
|
||||
|
||||
.macro _chacha_permute nrounds
|
||||
.set brot, 0
|
||||
.set drot, 0
|
||||
.rept \nrounds / 2
|
||||
_doubleround
|
||||
.endr
|
||||
.endm
|
||||
|
||||
.macro _chacha nrounds
|
||||
|
||||
.Lnext_block\@:
|
||||
// Stack: unused0-unused1 x10-x11 x0-x15 OUT IN LEN
|
||||
// Registers contain x0-x9,x12-x15.
|
||||
|
||||
// Do the core ChaCha permutation to update x0-x15.
|
||||
_chacha_permute \nrounds
|
||||
|
||||
add sp, #8
|
||||
// Stack: x10-x11 orig_x0-orig_x15 OUT IN LEN
|
||||
// Registers contain x0-x9,x12-x15.
|
||||
// x4-x7 are rotated by 'brot'; x12-x15 are rotated by 'drot'.
|
||||
|
||||
// Free up some registers (r8-r12,r14) by pushing (x8-x9,x12-x15).
|
||||
push {X8_X10, X9_X11, X12, X13, X14, X15}
|
||||
|
||||
// Load (OUT, IN, LEN).
|
||||
ldr r14, [sp, #96]
|
||||
ldr r12, [sp, #100]
|
||||
ldr r11, [sp, #104]
|
||||
|
||||
orr r10, r14, r12
|
||||
|
||||
// Use slow path if fewer than 64 bytes remain.
|
||||
cmp r11, #64
|
||||
blt .Lxor_slowpath\@
|
||||
|
||||
// Use slow path if IN and/or OUT isn't 4-byte aligned. Needed even on
|
||||
// ARMv6+, since ldmia and stmia (used below) still require alignment.
|
||||
tst r10, #3
|
||||
bne .Lxor_slowpath\@
|
||||
|
||||
// Fast path: XOR 64 bytes of aligned data.
|
||||
|
||||
// Stack: x8-x9 x12-x15 x10-x11 orig_x0-orig_x15 OUT IN LEN
|
||||
// Registers: r0-r7 are x0-x7; r8-r11 are free; r12 is IN; r14 is OUT.
|
||||
// x4-x7 are rotated by 'brot'; x12-x15 are rotated by 'drot'.
|
||||
|
||||
// x0-x3
|
||||
__ldrd r8, r9, sp, 32
|
||||
__ldrd r10, r11, sp, 40
|
||||
add X0, X0, r8
|
||||
add X1, X1, r9
|
||||
add X2, X2, r10
|
||||
add X3, X3, r11
|
||||
_le32_bswap_4x X0, X1, X2, X3, r8, r9, r10
|
||||
ldmia r12!, {r8-r11}
|
||||
eor X0, X0, r8
|
||||
eor X1, X1, r9
|
||||
eor X2, X2, r10
|
||||
eor X3, X3, r11
|
||||
stmia r14!, {X0-X3}
|
||||
|
||||
// x4-x7
|
||||
__ldrd r8, r9, sp, 48
|
||||
__ldrd r10, r11, sp, 56
|
||||
add X4, r8, X4, ror #brot
|
||||
add X5, r9, X5, ror #brot
|
||||
ldmia r12!, {X0-X3}
|
||||
add X6, r10, X6, ror #brot
|
||||
add X7, r11, X7, ror #brot
|
||||
_le32_bswap_4x X4, X5, X6, X7, r8, r9, r10
|
||||
eor X4, X4, X0
|
||||
eor X5, X5, X1
|
||||
eor X6, X6, X2
|
||||
eor X7, X7, X3
|
||||
stmia r14!, {X4-X7}
|
||||
|
||||
// x8-x15
|
||||
pop {r0-r7} // (x8-x9,x12-x15,x10-x11)
|
||||
__ldrd r8, r9, sp, 32
|
||||
__ldrd r10, r11, sp, 40
|
||||
add r0, r0, r8 // x8
|
||||
add r1, r1, r9 // x9
|
||||
add r6, r6, r10 // x10
|
||||
add r7, r7, r11 // x11
|
||||
_le32_bswap_4x r0, r1, r6, r7, r8, r9, r10
|
||||
ldmia r12!, {r8-r11}
|
||||
eor r0, r0, r8 // x8
|
||||
eor r1, r1, r9 // x9
|
||||
eor r6, r6, r10 // x10
|
||||
eor r7, r7, r11 // x11
|
||||
stmia r14!, {r0,r1,r6,r7}
|
||||
ldmia r12!, {r0,r1,r6,r7}
|
||||
__ldrd r8, r9, sp, 48
|
||||
__ldrd r10, r11, sp, 56
|
||||
add r2, r8, r2, ror #drot // x12
|
||||
add r3, r9, r3, ror #drot // x13
|
||||
add r4, r10, r4, ror #drot // x14
|
||||
add r5, r11, r5, ror #drot // x15
|
||||
_le32_bswap_4x r2, r3, r4, r5, r9, r10, r11
|
||||
ldr r9, [sp, #72] // load LEN
|
||||
eor r2, r2, r0 // x12
|
||||
eor r3, r3, r1 // x13
|
||||
eor r4, r4, r6 // x14
|
||||
eor r5, r5, r7 // x15
|
||||
subs r9, #64 // decrement and check LEN
|
||||
stmia r14!, {r2-r5}
|
||||
|
||||
beq .Ldone\@
|
||||
|
||||
.Lprepare_for_next_block\@:
|
||||
|
||||
// Stack: x0-x15 OUT IN LEN
|
||||
|
||||
// Increment block counter (x12)
|
||||
add r8, #1
|
||||
|
||||
// Store updated (OUT, IN, LEN)
|
||||
str r14, [sp, #64]
|
||||
str r12, [sp, #68]
|
||||
str r9, [sp, #72]
|
||||
|
||||
mov r14, sp
|
||||
|
||||
// Store updated block counter (x12)
|
||||
str r8, [sp, #48]
|
||||
|
||||
sub sp, #16
|
||||
|
||||
// Reload state and do next block
|
||||
ldmia r14!, {r0-r11} // load x0-x11
|
||||
__strd r10, r11, sp, 8 // store x10-x11 before state
|
||||
ldmia r14, {r10-r12,r14} // load x12-x15
|
||||
b .Lnext_block\@
|
||||
|
||||
.Lxor_slowpath\@:
|
||||
// Slow path: < 64 bytes remaining, or unaligned input or output buffer.
|
||||
// We handle it by storing the 64 bytes of keystream to the stack, then
|
||||
// XOR-ing the needed portion with the data.
|
||||
|
||||
// Allocate keystream buffer
|
||||
sub sp, #64
|
||||
mov r14, sp
|
||||
|
||||
// Stack: ks0-ks15 x8-x9 x12-x15 x10-x11 orig_x0-orig_x15 OUT IN LEN
|
||||
// Registers: r0-r7 are x0-x7; r8-r11 are free; r12 is IN; r14 is &ks0.
|
||||
// x4-x7 are rotated by 'brot'; x12-x15 are rotated by 'drot'.
|
||||
|
||||
// Save keystream for x0-x3
|
||||
__ldrd r8, r9, sp, 96
|
||||
__ldrd r10, r11, sp, 104
|
||||
add X0, X0, r8
|
||||
add X1, X1, r9
|
||||
add X2, X2, r10
|
||||
add X3, X3, r11
|
||||
_le32_bswap_4x X0, X1, X2, X3, r8, r9, r10
|
||||
stmia r14!, {X0-X3}
|
||||
|
||||
// Save keystream for x4-x7
|
||||
__ldrd r8, r9, sp, 112
|
||||
__ldrd r10, r11, sp, 120
|
||||
add X4, r8, X4, ror #brot
|
||||
add X5, r9, X5, ror #brot
|
||||
add X6, r10, X6, ror #brot
|
||||
add X7, r11, X7, ror #brot
|
||||
_le32_bswap_4x X4, X5, X6, X7, r8, r9, r10
|
||||
add r8, sp, #64
|
||||
stmia r14!, {X4-X7}
|
||||
|
||||
// Save keystream for x8-x15
|
||||
ldm r8, {r0-r7} // (x8-x9,x12-x15,x10-x11)
|
||||
__ldrd r8, r9, sp, 128
|
||||
__ldrd r10, r11, sp, 136
|
||||
add r0, r0, r8 // x8
|
||||
add r1, r1, r9 // x9
|
||||
add r6, r6, r10 // x10
|
||||
add r7, r7, r11 // x11
|
||||
_le32_bswap_4x r0, r1, r6, r7, r8, r9, r10
|
||||
stmia r14!, {r0,r1,r6,r7}
|
||||
__ldrd r8, r9, sp, 144
|
||||
__ldrd r10, r11, sp, 152
|
||||
add r2, r8, r2, ror #drot // x12
|
||||
add r3, r9, r3, ror #drot // x13
|
||||
add r4, r10, r4, ror #drot // x14
|
||||
add r5, r11, r5, ror #drot // x15
|
||||
_le32_bswap_4x r2, r3, r4, r5, r9, r10, r11
|
||||
stmia r14, {r2-r5}
|
||||
|
||||
// Stack: ks0-ks15 unused0-unused7 x0-x15 OUT IN LEN
|
||||
// Registers: r8 is block counter, r12 is IN.
|
||||
|
||||
ldr r9, [sp, #168] // LEN
|
||||
ldr r14, [sp, #160] // OUT
|
||||
cmp r9, #64
|
||||
mov r0, sp
|
||||
movle r1, r9
|
||||
movgt r1, #64
|
||||
// r1 is number of bytes to XOR, in range [1, 64]
|
||||
|
||||
.if __LINUX_ARM_ARCH__ < 6
|
||||
orr r2, r12, r14
|
||||
tst r2, #3 // IN or OUT misaligned?
|
||||
bne .Lxor_next_byte\@
|
||||
.endif
|
||||
|
||||
// XOR a word at a time
|
||||
.rept 16
|
||||
subs r1, #4
|
||||
blt .Lxor_words_done\@
|
||||
ldr r2, [r12], #4
|
||||
ldr r3, [r0], #4
|
||||
eor r2, r2, r3
|
||||
str r2, [r14], #4
|
||||
.endr
|
||||
b .Lxor_slowpath_done\@
|
||||
.Lxor_words_done\@:
|
||||
ands r1, r1, #3
|
||||
beq .Lxor_slowpath_done\@
|
||||
|
||||
// XOR a byte at a time
|
||||
.Lxor_next_byte\@:
|
||||
ldrb r2, [r12], #1
|
||||
ldrb r3, [r0], #1
|
||||
eor r2, r2, r3
|
||||
strb r2, [r14], #1
|
||||
subs r1, #1
|
||||
bne .Lxor_next_byte\@
|
||||
|
||||
.Lxor_slowpath_done\@:
|
||||
subs r9, #64
|
||||
add sp, #96
|
||||
bgt .Lprepare_for_next_block\@
|
||||
|
||||
.Ldone\@:
|
||||
.endm // _chacha
|
||||
|
||||
/*
|
||||
* void chacha_doarm(u8 *dst, const u8 *src, unsigned int bytes,
|
||||
* const u32 *state, int nrounds);
|
||||
*/
|
||||
ENTRY(chacha_doarm)
|
||||
cmp r2, #0 // len == 0?
|
||||
reteq lr
|
||||
|
||||
ldr ip, [sp]
|
||||
cmp ip, #12
|
||||
|
||||
push {r0-r2,r4-r11,lr}
|
||||
|
||||
// Push state x0-x15 onto stack.
|
||||
// Also store an extra copy of x10-x11 just before the state.
|
||||
|
||||
add X12, r3, #48
|
||||
ldm X12, {X12,X13,X14,X15}
|
||||
push {X12,X13,X14,X15}
|
||||
sub sp, sp, #64
|
||||
|
||||
__ldrd X8_X10, X9_X11, r3, 40
|
||||
__strd X8_X10, X9_X11, sp, 8
|
||||
__strd X8_X10, X9_X11, sp, 56
|
||||
ldm r3, {X0-X9_X11}
|
||||
__strd X0, X1, sp, 16
|
||||
__strd X2, X3, sp, 24
|
||||
__strd X4, X5, sp, 32
|
||||
__strd X6, X7, sp, 40
|
||||
__strd X8_X10, X9_X11, sp, 48
|
||||
|
||||
beq 1f
|
||||
_chacha 20
|
||||
|
||||
0: add sp, #76
|
||||
pop {r4-r11, pc}
|
||||
|
||||
1: _chacha 12
|
||||
b 0b
|
||||
ENDPROC(chacha_doarm)
|
||||
|
||||
/*
|
||||
* void hchacha_block_arm(const u32 state[16], u32 out[8], int nrounds);
|
||||
*/
|
||||
ENTRY(hchacha_block_arm)
|
||||
push {r1,r4-r11,lr}
|
||||
|
||||
cmp r2, #12 // ChaCha12 ?
|
||||
|
||||
mov r14, r0
|
||||
ldmia r14!, {r0-r11} // load x0-x11
|
||||
push {r10-r11} // store x10-x11 to stack
|
||||
ldm r14, {r10-r12,r14} // load x12-x15
|
||||
sub sp, #8
|
||||
|
||||
beq 1f
|
||||
_chacha_permute 20
|
||||
|
||||
// Skip over (unused0-unused1, x10-x11)
|
||||
0: add sp, #16
|
||||
|
||||
// Fix up rotations of x12-x15
|
||||
ror X12, X12, #drot
|
||||
ror X13, X13, #drot
|
||||
pop {r4} // load 'out'
|
||||
ror X14, X14, #drot
|
||||
ror X15, X15, #drot
|
||||
|
||||
// Store (x0-x3,x12-x15) to 'out'
|
||||
stm r4, {X0,X1,X2,X3,X12,X13,X14,X15}
|
||||
|
||||
pop {r4-r11,pc}
|
||||
|
||||
1: _chacha_permute 12
|
||||
b 0b
|
||||
ENDPROC(hchacha_block_arm)
|
||||
2062
arch/arm/crypto/curve25519-core.S
Normal file
2062
arch/arm/crypto/curve25519-core.S
Normal file
File diff suppressed because it is too large
Load diff
136
arch/arm/crypto/curve25519-glue.c
Normal file
136
arch/arm/crypto/curve25519-glue.c
Normal file
|
|
@ -0,0 +1,136 @@
|
|||
// SPDX-License-Identifier: GPL-2.0 OR MIT
|
||||
/*
|
||||
* Copyright (C) 2015-2019 Jason A. Donenfeld <Jason@zx2c4.com>. All Rights Reserved.
|
||||
*
|
||||
* Based on public domain code from Daniel J. Bernstein and Peter Schwabe. This
|
||||
* began from SUPERCOP's curve25519/neon2/scalarmult.s, but has subsequently been
|
||||
* manually reworked for use in kernel space.
|
||||
*/
|
||||
|
||||
#include <asm/hwcap.h>
|
||||
#include <asm/neon.h>
|
||||
#include <asm/simd.h>
|
||||
#include <crypto/internal/kpp.h>
|
||||
#include <crypto/internal/simd.h>
|
||||
#include <linux/types.h>
|
||||
#include <linux/module.h>
|
||||
#include <linux/init.h>
|
||||
#include <linux/jump_label.h>
|
||||
#include <linux/scatterlist.h>
|
||||
#include <crypto/curve25519.h>
|
||||
|
||||
asmlinkage void curve25519_neon(u8 mypublic[CURVE25519_KEY_SIZE],
|
||||
const u8 secret[CURVE25519_KEY_SIZE],
|
||||
const u8 basepoint[CURVE25519_KEY_SIZE]);
|
||||
|
||||
static __ro_after_init DEFINE_STATIC_KEY_FALSE(have_neon);
|
||||
|
||||
void curve25519_arch(u8 out[CURVE25519_KEY_SIZE],
|
||||
const u8 scalar[CURVE25519_KEY_SIZE],
|
||||
const u8 point[CURVE25519_KEY_SIZE])
|
||||
{
|
||||
if (static_branch_likely(&have_neon) && crypto_simd_usable()) {
|
||||
kernel_neon_begin();
|
||||
curve25519_neon(out, scalar, point);
|
||||
kernel_neon_end();
|
||||
} else {
|
||||
curve25519_generic(out, scalar, point);
|
||||
}
|
||||
}
|
||||
EXPORT_SYMBOL(curve25519_arch);
|
||||
|
||||
void curve25519_base_arch(u8 pub[CURVE25519_KEY_SIZE],
|
||||
const u8 secret[CURVE25519_KEY_SIZE])
|
||||
{
|
||||
return curve25519_arch(pub, secret, curve25519_base_point);
|
||||
}
|
||||
EXPORT_SYMBOL(curve25519_base_arch);
|
||||
|
||||
static int curve25519_set_secret(struct crypto_kpp *tfm, const void *buf,
|
||||
unsigned int len)
|
||||
{
|
||||
u8 *secret = kpp_tfm_ctx(tfm);
|
||||
|
||||
if (!len)
|
||||
curve25519_generate_secret(secret);
|
||||
else if (len == CURVE25519_KEY_SIZE &&
|
||||
crypto_memneq(buf, curve25519_null_point, CURVE25519_KEY_SIZE))
|
||||
memcpy(secret, buf, CURVE25519_KEY_SIZE);
|
||||
else
|
||||
return -EINVAL;
|
||||
return 0;
|
||||
}
|
||||
|
||||
static int curve25519_compute_value(struct kpp_request *req)
|
||||
{
|
||||
struct crypto_kpp *tfm = crypto_kpp_reqtfm(req);
|
||||
const u8 *secret = kpp_tfm_ctx(tfm);
|
||||
u8 public_key[CURVE25519_KEY_SIZE];
|
||||
u8 buf[CURVE25519_KEY_SIZE];
|
||||
int copied, nbytes;
|
||||
u8 const *bp;
|
||||
|
||||
if (req->src) {
|
||||
copied = sg_copy_to_buffer(req->src,
|
||||
sg_nents_for_len(req->src,
|
||||
CURVE25519_KEY_SIZE),
|
||||
public_key, CURVE25519_KEY_SIZE);
|
||||
if (copied != CURVE25519_KEY_SIZE)
|
||||
return -EINVAL;
|
||||
bp = public_key;
|
||||
} else {
|
||||
bp = curve25519_base_point;
|
||||
}
|
||||
|
||||
curve25519_arch(buf, secret, bp);
|
||||
|
||||
/* might want less than we've got */
|
||||
nbytes = min_t(size_t, CURVE25519_KEY_SIZE, req->dst_len);
|
||||
copied = sg_copy_from_buffer(req->dst, sg_nents_for_len(req->dst,
|
||||
nbytes),
|
||||
buf, nbytes);
|
||||
if (copied != nbytes)
|
||||
return -EINVAL;
|
||||
return 0;
|
||||
}
|
||||
|
||||
static unsigned int curve25519_max_size(struct crypto_kpp *tfm)
|
||||
{
|
||||
return CURVE25519_KEY_SIZE;
|
||||
}
|
||||
|
||||
static struct kpp_alg curve25519_alg = {
|
||||
.base.cra_name = "curve25519",
|
||||
.base.cra_driver_name = "curve25519-neon",
|
||||
.base.cra_priority = 200,
|
||||
.base.cra_module = THIS_MODULE,
|
||||
.base.cra_ctxsize = CURVE25519_KEY_SIZE,
|
||||
|
||||
.set_secret = curve25519_set_secret,
|
||||
.generate_public_key = curve25519_compute_value,
|
||||
.compute_shared_secret = curve25519_compute_value,
|
||||
.max_size = curve25519_max_size,
|
||||
};
|
||||
|
||||
static int __init mod_init(void)
|
||||
{
|
||||
if (elf_hwcap & HWCAP_NEON) {
|
||||
static_branch_enable(&have_neon);
|
||||
return IS_REACHABLE(CONFIG_CRYPTO_KPP) ?
|
||||
crypto_register_kpp(&curve25519_alg) : 0;
|
||||
}
|
||||
return 0;
|
||||
}
|
||||
|
||||
static void __exit mod_exit(void)
|
||||
{
|
||||
if (IS_REACHABLE(CONFIG_CRYPTO_KPP) && elf_hwcap & HWCAP_NEON)
|
||||
crypto_unregister_kpp(&curve25519_alg);
|
||||
}
|
||||
|
||||
module_init(mod_init);
|
||||
module_exit(mod_exit);
|
||||
|
||||
MODULE_ALIAS_CRYPTO("curve25519");
|
||||
MODULE_ALIAS_CRYPTO("curve25519-neon");
|
||||
MODULE_LICENSE("GPL v2");
|
||||
1236
arch/arm/crypto/poly1305-armv4.pl
Normal file
1236
arch/arm/crypto/poly1305-armv4.pl
Normal file
File diff suppressed because it is too large
Load diff
1158
arch/arm/crypto/poly1305-core.S_shipped
Normal file
1158
arch/arm/crypto/poly1305-core.S_shipped
Normal file
File diff suppressed because it is too large
Load diff
273
arch/arm/crypto/poly1305-glue.c
Normal file
273
arch/arm/crypto/poly1305-glue.c
Normal file
|
|
@ -0,0 +1,273 @@
|
|||
// SPDX-License-Identifier: GPL-2.0
|
||||
/*
|
||||
* OpenSSL/Cryptogams accelerated Poly1305 transform for ARM
|
||||
*
|
||||
* Copyright (C) 2019 Linaro Ltd. <ard.biesheuvel@linaro.org>
|
||||
*/
|
||||
|
||||
#include <asm/hwcap.h>
|
||||
#include <asm/neon.h>
|
||||
#include <asm/simd.h>
|
||||
#include <asm/unaligned.h>
|
||||
#include <crypto/algapi.h>
|
||||
#include <crypto/internal/hash.h>
|
||||
#include <crypto/internal/poly1305.h>
|
||||
#include <crypto/internal/simd.h>
|
||||
#include <linux/cpufeature.h>
|
||||
#include <linux/crypto.h>
|
||||
#include <linux/jump_label.h>
|
||||
#include <linux/module.h>
|
||||
|
||||
void poly1305_init_arm(void *state, const u8 *key);
|
||||
void poly1305_blocks_arm(void *state, const u8 *src, u32 len, u32 hibit);
|
||||
void poly1305_blocks_neon(void *state, const u8 *src, u32 len, u32 hibit);
|
||||
void poly1305_emit_arm(void *state, u8 *digest, const u32 *nonce);
|
||||
|
||||
void __weak poly1305_blocks_neon(void *state, const u8 *src, u32 len, u32 hibit)
|
||||
{
|
||||
}
|
||||
|
||||
static __ro_after_init DEFINE_STATIC_KEY_FALSE(have_neon);
|
||||
|
||||
void poly1305_init_arch(struct poly1305_desc_ctx *dctx, const u8 key[POLY1305_KEY_SIZE])
|
||||
{
|
||||
poly1305_init_arm(&dctx->h, key);
|
||||
dctx->s[0] = get_unaligned_le32(key + 16);
|
||||
dctx->s[1] = get_unaligned_le32(key + 20);
|
||||
dctx->s[2] = get_unaligned_le32(key + 24);
|
||||
dctx->s[3] = get_unaligned_le32(key + 28);
|
||||
dctx->buflen = 0;
|
||||
}
|
||||
EXPORT_SYMBOL(poly1305_init_arch);
|
||||
|
||||
static int arm_poly1305_init(struct shash_desc *desc)
|
||||
{
|
||||
struct poly1305_desc_ctx *dctx = shash_desc_ctx(desc);
|
||||
|
||||
dctx->buflen = 0;
|
||||
dctx->rset = 0;
|
||||
dctx->sset = false;
|
||||
|
||||
return 0;
|
||||
}
|
||||
|
||||
static void arm_poly1305_blocks(struct poly1305_desc_ctx *dctx, const u8 *src,
|
||||
u32 len, u32 hibit, bool do_neon)
|
||||
{
|
||||
if (unlikely(!dctx->sset)) {
|
||||
if (!dctx->rset) {
|
||||
poly1305_init_arm(&dctx->h, src);
|
||||
src += POLY1305_BLOCK_SIZE;
|
||||
len -= POLY1305_BLOCK_SIZE;
|
||||
dctx->rset = 1;
|
||||
}
|
||||
if (len >= POLY1305_BLOCK_SIZE) {
|
||||
dctx->s[0] = get_unaligned_le32(src + 0);
|
||||
dctx->s[1] = get_unaligned_le32(src + 4);
|
||||
dctx->s[2] = get_unaligned_le32(src + 8);
|
||||
dctx->s[3] = get_unaligned_le32(src + 12);
|
||||
src += POLY1305_BLOCK_SIZE;
|
||||
len -= POLY1305_BLOCK_SIZE;
|
||||
dctx->sset = true;
|
||||
}
|
||||
if (len < POLY1305_BLOCK_SIZE)
|
||||
return;
|
||||
}
|
||||
|
||||
len &= ~(POLY1305_BLOCK_SIZE - 1);
|
||||
|
||||
if (static_branch_likely(&have_neon) && likely(do_neon))
|
||||
poly1305_blocks_neon(&dctx->h, src, len, hibit);
|
||||
else
|
||||
poly1305_blocks_arm(&dctx->h, src, len, hibit);
|
||||
}
|
||||
|
||||
static void arm_poly1305_do_update(struct poly1305_desc_ctx *dctx,
|
||||
const u8 *src, u32 len, bool do_neon)
|
||||
{
|
||||
if (unlikely(dctx->buflen)) {
|
||||
u32 bytes = min(len, POLY1305_BLOCK_SIZE - dctx->buflen);
|
||||
|
||||
memcpy(dctx->buf + dctx->buflen, src, bytes);
|
||||
src += bytes;
|
||||
len -= bytes;
|
||||
dctx->buflen += bytes;
|
||||
|
||||
if (dctx->buflen == POLY1305_BLOCK_SIZE) {
|
||||
arm_poly1305_blocks(dctx, dctx->buf,
|
||||
POLY1305_BLOCK_SIZE, 1, false);
|
||||
dctx->buflen = 0;
|
||||
}
|
||||
}
|
||||
|
||||
if (likely(len >= POLY1305_BLOCK_SIZE)) {
|
||||
arm_poly1305_blocks(dctx, src, len, 1, do_neon);
|
||||
src += round_down(len, POLY1305_BLOCK_SIZE);
|
||||
len %= POLY1305_BLOCK_SIZE;
|
||||
}
|
||||
|
||||
if (unlikely(len)) {
|
||||
dctx->buflen = len;
|
||||
memcpy(dctx->buf, src, len);
|
||||
}
|
||||
}
|
||||
|
||||
static int arm_poly1305_update(struct shash_desc *desc,
|
||||
const u8 *src, unsigned int srclen)
|
||||
{
|
||||
struct poly1305_desc_ctx *dctx = shash_desc_ctx(desc);
|
||||
|
||||
arm_poly1305_do_update(dctx, src, srclen, false);
|
||||
return 0;
|
||||
}
|
||||
|
||||
static int __maybe_unused arm_poly1305_update_neon(struct shash_desc *desc,
|
||||
const u8 *src,
|
||||
unsigned int srclen)
|
||||
{
|
||||
struct poly1305_desc_ctx *dctx = shash_desc_ctx(desc);
|
||||
bool do_neon = crypto_simd_usable() && srclen > 128;
|
||||
|
||||
if (static_branch_likely(&have_neon) && do_neon)
|
||||
kernel_neon_begin();
|
||||
arm_poly1305_do_update(dctx, src, srclen, do_neon);
|
||||
if (static_branch_likely(&have_neon) && do_neon)
|
||||
kernel_neon_end();
|
||||
return 0;
|
||||
}
|
||||
|
||||
void poly1305_update_arch(struct poly1305_desc_ctx *dctx, const u8 *src,
|
||||
unsigned int nbytes)
|
||||
{
|
||||
bool do_neon = IS_ENABLED(CONFIG_KERNEL_MODE_NEON) &&
|
||||
crypto_simd_usable();
|
||||
|
||||
if (unlikely(dctx->buflen)) {
|
||||
u32 bytes = min(nbytes, POLY1305_BLOCK_SIZE - dctx->buflen);
|
||||
|
||||
memcpy(dctx->buf + dctx->buflen, src, bytes);
|
||||
src += bytes;
|
||||
nbytes -= bytes;
|
||||
dctx->buflen += bytes;
|
||||
|
||||
if (dctx->buflen == POLY1305_BLOCK_SIZE) {
|
||||
poly1305_blocks_arm(&dctx->h, dctx->buf,
|
||||
POLY1305_BLOCK_SIZE, 1);
|
||||
dctx->buflen = 0;
|
||||
}
|
||||
}
|
||||
|
||||
if (likely(nbytes >= POLY1305_BLOCK_SIZE)) {
|
||||
unsigned int len = round_down(nbytes, POLY1305_BLOCK_SIZE);
|
||||
|
||||
if (static_branch_likely(&have_neon) && do_neon) {
|
||||
do {
|
||||
unsigned int todo = min_t(unsigned int, len, SZ_4K);
|
||||
|
||||
kernel_neon_begin();
|
||||
poly1305_blocks_neon(&dctx->h, src, todo, 1);
|
||||
kernel_neon_end();
|
||||
|
||||
len -= todo;
|
||||
src += todo;
|
||||
} while (len);
|
||||
} else {
|
||||
poly1305_blocks_arm(&dctx->h, src, len, 1);
|
||||
src += len;
|
||||
}
|
||||
nbytes %= POLY1305_BLOCK_SIZE;
|
||||
}
|
||||
|
||||
if (unlikely(nbytes)) {
|
||||
dctx->buflen = nbytes;
|
||||
memcpy(dctx->buf, src, nbytes);
|
||||
}
|
||||
}
|
||||
EXPORT_SYMBOL(poly1305_update_arch);
|
||||
|
||||
void poly1305_final_arch(struct poly1305_desc_ctx *dctx, u8 *dst)
|
||||
{
|
||||
if (unlikely(dctx->buflen)) {
|
||||
dctx->buf[dctx->buflen++] = 1;
|
||||
memset(dctx->buf + dctx->buflen, 0,
|
||||
POLY1305_BLOCK_SIZE - dctx->buflen);
|
||||
poly1305_blocks_arm(&dctx->h, dctx->buf, POLY1305_BLOCK_SIZE, 0);
|
||||
}
|
||||
|
||||
poly1305_emit_arm(&dctx->h, dst, dctx->s);
|
||||
*dctx = (struct poly1305_desc_ctx){};
|
||||
}
|
||||
EXPORT_SYMBOL(poly1305_final_arch);
|
||||
|
||||
static int arm_poly1305_final(struct shash_desc *desc, u8 *dst)
|
||||
{
|
||||
struct poly1305_desc_ctx *dctx = shash_desc_ctx(desc);
|
||||
|
||||
if (unlikely(!dctx->sset))
|
||||
return -ENOKEY;
|
||||
|
||||
poly1305_final_arch(dctx, dst);
|
||||
return 0;
|
||||
}
|
||||
|
||||
static struct shash_alg arm_poly1305_algs[] = {{
|
||||
.init = arm_poly1305_init,
|
||||
.update = arm_poly1305_update,
|
||||
.final = arm_poly1305_final,
|
||||
.digestsize = POLY1305_DIGEST_SIZE,
|
||||
.descsize = sizeof(struct poly1305_desc_ctx),
|
||||
|
||||
.base.cra_name = "poly1305",
|
||||
.base.cra_driver_name = "poly1305-arm",
|
||||
.base.cra_priority = 150,
|
||||
.base.cra_blocksize = POLY1305_BLOCK_SIZE,
|
||||
.base.cra_module = THIS_MODULE,
|
||||
#ifdef CONFIG_KERNEL_MODE_NEON
|
||||
}, {
|
||||
.init = arm_poly1305_init,
|
||||
.update = arm_poly1305_update_neon,
|
||||
.final = arm_poly1305_final,
|
||||
.digestsize = POLY1305_DIGEST_SIZE,
|
||||
.descsize = sizeof(struct poly1305_desc_ctx),
|
||||
|
||||
.base.cra_name = "poly1305",
|
||||
.base.cra_driver_name = "poly1305-neon",
|
||||
.base.cra_priority = 200,
|
||||
.base.cra_blocksize = POLY1305_BLOCK_SIZE,
|
||||
.base.cra_module = THIS_MODULE,
|
||||
#endif
|
||||
}};
|
||||
|
||||
static int __init arm_poly1305_mod_init(void)
|
||||
{
|
||||
if (IS_ENABLED(CONFIG_KERNEL_MODE_NEON) &&
|
||||
(elf_hwcap & HWCAP_NEON))
|
||||
static_branch_enable(&have_neon);
|
||||
else if (IS_REACHABLE(CONFIG_CRYPTO_HASH))
|
||||
/* register only the first entry */
|
||||
return crypto_register_shash(&arm_poly1305_algs[0]);
|
||||
|
||||
return IS_REACHABLE(CONFIG_CRYPTO_HASH) ?
|
||||
crypto_register_shashes(arm_poly1305_algs,
|
||||
ARRAY_SIZE(arm_poly1305_algs)) : 0;
|
||||
}
|
||||
|
||||
static void __exit arm_poly1305_mod_exit(void)
|
||||
{
|
||||
if (!IS_REACHABLE(CONFIG_CRYPTO_HASH))
|
||||
return;
|
||||
if (!static_branch_likely(&have_neon)) {
|
||||
crypto_unregister_shash(&arm_poly1305_algs[0]);
|
||||
return;
|
||||
}
|
||||
crypto_unregister_shashes(arm_poly1305_algs,
|
||||
ARRAY_SIZE(arm_poly1305_algs));
|
||||
}
|
||||
|
||||
module_init(arm_poly1305_mod_init);
|
||||
module_exit(arm_poly1305_mod_exit);
|
||||
|
||||
MODULE_LICENSE("GPL v2");
|
||||
MODULE_ALIAS_CRYPTO("poly1305");
|
||||
MODULE_ALIAS_CRYPTO("poly1305-arm");
|
||||
MODULE_ALIAS_CRYPTO("poly1305-neon");
|
||||
|
|
@ -83,7 +83,6 @@ CONFIG_ARM_SCMI_PROTOCOL=y
|
|||
CONFIG_ARM_SCPI_PROTOCOL=y
|
||||
# CONFIG_ARM_SCPI_POWER_DOMAIN is not set
|
||||
# CONFIG_EFI_ARMSTUB_DTB_LOADER is not set
|
||||
CONFIG_ARM64_CRYPTO=y
|
||||
CONFIG_CRYPTO_SHA2_ARM64_CE=y
|
||||
CONFIG_CRYPTO_AES_ARM64_CE_BLK=y
|
||||
CONFIG_KPROBES=y
|
||||
|
|
@ -275,6 +274,7 @@ CONFIG_DM_VERITY_FEC=y
|
|||
CONFIG_DM_BOW=y
|
||||
CONFIG_NETDEVICES=y
|
||||
CONFIG_DUMMY=y
|
||||
CONFIG_WIREGUARD=y
|
||||
CONFIG_TUN=y
|
||||
CONFIG_VETH=y
|
||||
# CONFIG_ETHERNET is not set
|
||||
|
|
|
|||
1
arch/arm64/crypto/.gitignore
vendored
1
arch/arm64/crypto/.gitignore
vendored
|
|
@ -1,2 +1,3 @@
|
|||
sha256-core.S
|
||||
sha512-core.S
|
||||
poly1305-core.S
|
||||
|
|
|
|||
|
|
@ -104,7 +104,14 @@ config CRYPTO_CHACHA20_NEON
|
|||
tristate "ChaCha20, XChaCha20, and XChaCha12 stream ciphers using NEON instructions"
|
||||
depends on KERNEL_MODE_NEON
|
||||
select CRYPTO_BLKCIPHER
|
||||
select CRYPTO_CHACHA20
|
||||
select CRYPTO_LIB_CHACHA_GENERIC
|
||||
select CRYPTO_ARCH_HAVE_LIB_CHACHA
|
||||
|
||||
config CRYPTO_POLY1305_NEON
|
||||
tristate "Poly1305 hash function using scalar or NEON instructions"
|
||||
depends on KERNEL_MODE_NEON
|
||||
select CRYPTO_HASH
|
||||
select CRYPTO_ARCH_HAVE_LIB_POLY1305
|
||||
|
||||
config CRYPTO_NHPOLY1305_NEON
|
||||
tristate "NHPoly1305 hash function using NEON instructions (for Adiantum)"
|
||||
|
|
|
|||
|
|
@ -50,6 +50,10 @@ sha512-arm64-y := sha512-glue.o sha512-core.o
|
|||
obj-$(CONFIG_CRYPTO_CHACHA20_NEON) += chacha-neon.o
|
||||
chacha-neon-y := chacha-neon-core.o chacha-neon-glue.o
|
||||
|
||||
obj-$(CONFIG_CRYPTO_POLY1305_NEON) += poly1305-neon.o
|
||||
poly1305-neon-y := poly1305-core.o poly1305-glue.o
|
||||
AFLAGS_poly1305-core.o += -Dpoly1305_init=poly1305_init_arm64
|
||||
|
||||
obj-$(CONFIG_CRYPTO_NHPOLY1305_NEON) += nhpoly1305-neon.o
|
||||
nhpoly1305-neon-y := nh-neon-core.o nhpoly1305-neon-glue.o
|
||||
|
||||
|
|
@ -68,11 +72,15 @@ ifdef REGENERATE_ARM64_CRYPTO
|
|||
quiet_cmd_perlasm = PERLASM $@
|
||||
cmd_perlasm = $(PERL) $(<) void $(@)
|
||||
|
||||
$(src)/poly1305-core.S_shipped: $(src)/poly1305-armv8.pl
|
||||
$(call cmd,perlasm)
|
||||
|
||||
$(src)/sha256-core.S_shipped: $(src)/sha512-armv8.pl
|
||||
$(call cmd,perlasm)
|
||||
|
||||
$(src)/sha512-core.S_shipped: $(src)/sha512-armv8.pl
|
||||
$(call cmd,perlasm)
|
||||
|
||||
endif
|
||||
|
||||
clean-files += sha256-core.S sha512-core.S
|
||||
clean-files += poly1305-core.S sha256-core.S sha512-core.S
|
||||
|
|
|
|||
|
|
@ -1,5 +1,5 @@
|
|||
/*
|
||||
* ARM NEON accelerated ChaCha and XChaCha stream ciphers,
|
||||
* ARM NEON and scalar accelerated ChaCha and XChaCha stream ciphers,
|
||||
* including ChaCha20 (RFC7539)
|
||||
*
|
||||
* Copyright (C) 2016 - 2017 Linaro, Ltd. <ard.biesheuvel@linaro.org>
|
||||
|
|
@ -20,9 +20,10 @@
|
|||
*/
|
||||
|
||||
#include <crypto/algapi.h>
|
||||
#include <crypto/chacha.h>
|
||||
#include <crypto/internal/chacha.h>
|
||||
#include <crypto/internal/simd.h>
|
||||
#include <crypto/internal/skcipher.h>
|
||||
#include <linux/jump_label.h>
|
||||
#include <linux/kernel.h>
|
||||
#include <linux/module.h>
|
||||
|
||||
|
|
@ -36,6 +37,8 @@ asmlinkage void chacha_4block_xor_neon(u32 *state, u8 *dst, const u8 *src,
|
|||
int nrounds, int bytes);
|
||||
asmlinkage void hchacha_block_neon(const u32 *state, u32 *out, int nrounds);
|
||||
|
||||
static __ro_after_init DEFINE_STATIC_KEY_FALSE(have_neon);
|
||||
|
||||
static void chacha_doneon(u32 *state, u8 *dst, const u8 *src,
|
||||
int bytes, int nrounds)
|
||||
{
|
||||
|
|
@ -52,13 +55,52 @@ static void chacha_doneon(u32 *state, u8 *dst, const u8 *src,
|
|||
break;
|
||||
}
|
||||
chacha_4block_xor_neon(state, dst, src, nrounds, l);
|
||||
bytes -= CHACHA_BLOCK_SIZE * 5;
|
||||
src += CHACHA_BLOCK_SIZE * 5;
|
||||
dst += CHACHA_BLOCK_SIZE * 5;
|
||||
state[12] += 5;
|
||||
bytes -= l;
|
||||
src += l;
|
||||
dst += l;
|
||||
state[12] += DIV_ROUND_UP(l, CHACHA_BLOCK_SIZE);
|
||||
}
|
||||
}
|
||||
|
||||
void hchacha_block_arch(const u32 *state, u32 *stream, int nrounds)
|
||||
{
|
||||
if (!static_branch_likely(&have_neon) || !crypto_simd_usable()) {
|
||||
hchacha_block_generic(state, stream, nrounds);
|
||||
} else {
|
||||
kernel_neon_begin();
|
||||
hchacha_block_neon(state, stream, nrounds);
|
||||
kernel_neon_end();
|
||||
}
|
||||
}
|
||||
EXPORT_SYMBOL(hchacha_block_arch);
|
||||
|
||||
void chacha_init_arch(u32 *state, const u32 *key, const u8 *iv)
|
||||
{
|
||||
chacha_init_generic(state, key, iv);
|
||||
}
|
||||
EXPORT_SYMBOL(chacha_init_arch);
|
||||
|
||||
void chacha_crypt_arch(u32 *state, u8 *dst, const u8 *src, unsigned int bytes,
|
||||
int nrounds)
|
||||
{
|
||||
if (!static_branch_likely(&have_neon) || bytes <= CHACHA_BLOCK_SIZE ||
|
||||
!crypto_simd_usable())
|
||||
return chacha_crypt_generic(state, dst, src, bytes, nrounds);
|
||||
|
||||
do {
|
||||
unsigned int todo = min_t(unsigned int, bytes, SZ_4K);
|
||||
|
||||
kernel_neon_begin();
|
||||
chacha_doneon(state, dst, src, todo, nrounds);
|
||||
kernel_neon_end();
|
||||
|
||||
bytes -= todo;
|
||||
src += todo;
|
||||
dst += todo;
|
||||
} while (bytes);
|
||||
}
|
||||
EXPORT_SYMBOL(chacha_crypt_arch);
|
||||
|
||||
static int chacha_neon_stream_xor(struct skcipher_request *req,
|
||||
const struct chacha_ctx *ctx, const u8 *iv)
|
||||
{
|
||||
|
|
@ -68,7 +110,7 @@ static int chacha_neon_stream_xor(struct skcipher_request *req,
|
|||
|
||||
err = skcipher_walk_virt(&walk, req, false);
|
||||
|
||||
crypto_chacha_init(state, ctx, iv);
|
||||
chacha_init_generic(state, ctx->key, iv);
|
||||
|
||||
while (walk.nbytes > 0) {
|
||||
unsigned int nbytes = walk.nbytes;
|
||||
|
|
@ -76,10 +118,17 @@ static int chacha_neon_stream_xor(struct skcipher_request *req,
|
|||
if (nbytes < walk.total)
|
||||
nbytes = rounddown(nbytes, walk.stride);
|
||||
|
||||
kernel_neon_begin();
|
||||
chacha_doneon(state, walk.dst.virt.addr, walk.src.virt.addr,
|
||||
nbytes, ctx->nrounds);
|
||||
kernel_neon_end();
|
||||
if (!static_branch_likely(&have_neon) ||
|
||||
!crypto_simd_usable()) {
|
||||
chacha_crypt_generic(state, walk.dst.virt.addr,
|
||||
walk.src.virt.addr, nbytes,
|
||||
ctx->nrounds);
|
||||
} else {
|
||||
kernel_neon_begin();
|
||||
chacha_doneon(state, walk.dst.virt.addr,
|
||||
walk.src.virt.addr, nbytes, ctx->nrounds);
|
||||
kernel_neon_end();
|
||||
}
|
||||
err = skcipher_walk_done(&walk, walk.nbytes - nbytes);
|
||||
}
|
||||
|
||||
|
|
@ -91,9 +140,6 @@ static int chacha_neon(struct skcipher_request *req)
|
|||
struct crypto_skcipher *tfm = crypto_skcipher_reqtfm(req);
|
||||
struct chacha_ctx *ctx = crypto_skcipher_ctx(tfm);
|
||||
|
||||
if (req->cryptlen <= CHACHA_BLOCK_SIZE || !crypto_simd_usable())
|
||||
return crypto_chacha_crypt(req);
|
||||
|
||||
return chacha_neon_stream_xor(req, ctx, req->iv);
|
||||
}
|
||||
|
||||
|
|
@ -105,14 +151,8 @@ static int xchacha_neon(struct skcipher_request *req)
|
|||
u32 state[16];
|
||||
u8 real_iv[16];
|
||||
|
||||
if (req->cryptlen <= CHACHA_BLOCK_SIZE || !crypto_simd_usable())
|
||||
return crypto_xchacha_crypt(req);
|
||||
|
||||
crypto_chacha_init(state, ctx, req->iv);
|
||||
|
||||
kernel_neon_begin();
|
||||
hchacha_block_neon(state, subctx.key, ctx->nrounds);
|
||||
kernel_neon_end();
|
||||
chacha_init_generic(state, ctx->key, req->iv);
|
||||
hchacha_block_arch(state, subctx.key, ctx->nrounds);
|
||||
subctx.nrounds = ctx->nrounds;
|
||||
|
||||
memcpy(&real_iv[0], req->iv + 24, 8);
|
||||
|
|
@ -134,7 +174,7 @@ static struct skcipher_alg algs[] = {
|
|||
.ivsize = CHACHA_IV_SIZE,
|
||||
.chunksize = CHACHA_BLOCK_SIZE,
|
||||
.walksize = 5 * CHACHA_BLOCK_SIZE,
|
||||
.setkey = crypto_chacha20_setkey,
|
||||
.setkey = chacha20_setkey,
|
||||
.encrypt = chacha_neon,
|
||||
.decrypt = chacha_neon,
|
||||
}, {
|
||||
|
|
@ -150,7 +190,7 @@ static struct skcipher_alg algs[] = {
|
|||
.ivsize = XCHACHA_IV_SIZE,
|
||||
.chunksize = CHACHA_BLOCK_SIZE,
|
||||
.walksize = 5 * CHACHA_BLOCK_SIZE,
|
||||
.setkey = crypto_chacha20_setkey,
|
||||
.setkey = chacha20_setkey,
|
||||
.encrypt = xchacha_neon,
|
||||
.decrypt = xchacha_neon,
|
||||
}, {
|
||||
|
|
@ -166,7 +206,7 @@ static struct skcipher_alg algs[] = {
|
|||
.ivsize = XCHACHA_IV_SIZE,
|
||||
.chunksize = CHACHA_BLOCK_SIZE,
|
||||
.walksize = 5 * CHACHA_BLOCK_SIZE,
|
||||
.setkey = crypto_chacha12_setkey,
|
||||
.setkey = chacha12_setkey,
|
||||
.encrypt = xchacha_neon,
|
||||
.decrypt = xchacha_neon,
|
||||
}
|
||||
|
|
@ -175,14 +215,18 @@ static struct skcipher_alg algs[] = {
|
|||
static int __init chacha_simd_mod_init(void)
|
||||
{
|
||||
if (!cpu_have_named_feature(ASIMD))
|
||||
return -ENODEV;
|
||||
return 0;
|
||||
|
||||
return crypto_register_skciphers(algs, ARRAY_SIZE(algs));
|
||||
static_branch_enable(&have_neon);
|
||||
|
||||
return IS_REACHABLE(CONFIG_CRYPTO_BLKCIPHER) ?
|
||||
crypto_register_skciphers(algs, ARRAY_SIZE(algs)) : 0;
|
||||
}
|
||||
|
||||
static void __exit chacha_simd_mod_fini(void)
|
||||
{
|
||||
crypto_unregister_skciphers(algs, ARRAY_SIZE(algs));
|
||||
if (IS_REACHABLE(CONFIG_CRYPTO_BLKCIPHER) && cpu_have_named_feature(ASIMD))
|
||||
crypto_unregister_skciphers(algs, ARRAY_SIZE(algs));
|
||||
}
|
||||
|
||||
module_init(chacha_simd_mod_init);
|
||||
|
|
|
|||
913
arch/arm64/crypto/poly1305-armv8.pl
Normal file
913
arch/arm64/crypto/poly1305-armv8.pl
Normal file
|
|
@ -0,0 +1,913 @@
|
|||
#!/usr/bin/env perl
|
||||
# SPDX-License-Identifier: GPL-1.0+ OR BSD-3-Clause
|
||||
#
|
||||
# ====================================================================
|
||||
# Written by Andy Polyakov, @dot-asm, initially for the OpenSSL
|
||||
# project.
|
||||
# ====================================================================
|
||||
#
|
||||
# This module implements Poly1305 hash for ARMv8.
|
||||
#
|
||||
# June 2015
|
||||
#
|
||||
# Numbers are cycles per processed byte with poly1305_blocks alone.
|
||||
#
|
||||
# IALU/gcc-4.9 NEON
|
||||
#
|
||||
# Apple A7 1.86/+5% 0.72
|
||||
# Cortex-A53 2.69/+58% 1.47
|
||||
# Cortex-A57 2.70/+7% 1.14
|
||||
# Denver 1.64/+50% 1.18(*)
|
||||
# X-Gene 2.13/+68% 2.27
|
||||
# Mongoose 1.77/+75% 1.12
|
||||
# Kryo 2.70/+55% 1.13
|
||||
# ThunderX2 1.17/+95% 1.36
|
||||
#
|
||||
# (*) estimate based on resources availability is less than 1.0,
|
||||
# i.e. measured result is worse than expected, presumably binary
|
||||
# translator is not almighty;
|
||||
|
||||
$flavour=shift;
|
||||
$output=shift;
|
||||
|
||||
if ($flavour && $flavour ne "void") {
|
||||
$0 =~ m/(.*[\/\\])[^\/\\]+$/; $dir=$1;
|
||||
( $xlate="${dir}arm-xlate.pl" and -f $xlate ) or
|
||||
( $xlate="${dir}../../perlasm/arm-xlate.pl" and -f $xlate) or
|
||||
die "can't locate arm-xlate.pl";
|
||||
|
||||
open STDOUT,"| \"$^X\" $xlate $flavour $output";
|
||||
} else {
|
||||
open STDOUT,">$output";
|
||||
}
|
||||
|
||||
my ($ctx,$inp,$len,$padbit) = map("x$_",(0..3));
|
||||
my ($mac,$nonce)=($inp,$len);
|
||||
|
||||
my ($h0,$h1,$h2,$r0,$r1,$s1,$t0,$t1,$d0,$d1,$d2) = map("x$_",(4..14));
|
||||
|
||||
$code.=<<___;
|
||||
#ifndef __KERNEL__
|
||||
# include "arm_arch.h"
|
||||
.extern OPENSSL_armcap_P
|
||||
#endif
|
||||
|
||||
.text
|
||||
|
||||
// forward "declarations" are required for Apple
|
||||
.globl poly1305_blocks
|
||||
.globl poly1305_emit
|
||||
|
||||
.globl poly1305_init
|
||||
.type poly1305_init,%function
|
||||
.align 5
|
||||
poly1305_init:
|
||||
cmp $inp,xzr
|
||||
stp xzr,xzr,[$ctx] // zero hash value
|
||||
stp xzr,xzr,[$ctx,#16] // [along with is_base2_26]
|
||||
|
||||
csel x0,xzr,x0,eq
|
||||
b.eq .Lno_key
|
||||
|
||||
#ifndef __KERNEL__
|
||||
adrp x17,OPENSSL_armcap_P
|
||||
ldr w17,[x17,#:lo12:OPENSSL_armcap_P]
|
||||
#endif
|
||||
|
||||
ldp $r0,$r1,[$inp] // load key
|
||||
mov $s1,#0xfffffffc0fffffff
|
||||
movk $s1,#0x0fff,lsl#48
|
||||
#ifdef __AARCH64EB__
|
||||
rev $r0,$r0 // flip bytes
|
||||
rev $r1,$r1
|
||||
#endif
|
||||
and $r0,$r0,$s1 // &=0ffffffc0fffffff
|
||||
and $s1,$s1,#-4
|
||||
and $r1,$r1,$s1 // &=0ffffffc0ffffffc
|
||||
mov w#$s1,#-1
|
||||
stp $r0,$r1,[$ctx,#32] // save key value
|
||||
str w#$s1,[$ctx,#48] // impossible key power value
|
||||
|
||||
#ifndef __KERNEL__
|
||||
tst w17,#ARMV7_NEON
|
||||
|
||||
adr $d0,.Lpoly1305_blocks
|
||||
adr $r0,.Lpoly1305_blocks_neon
|
||||
adr $d1,.Lpoly1305_emit
|
||||
|
||||
csel $d0,$d0,$r0,eq
|
||||
|
||||
# ifdef __ILP32__
|
||||
stp w#$d0,w#$d1,[$len]
|
||||
# else
|
||||
stp $d0,$d1,[$len]
|
||||
# endif
|
||||
#endif
|
||||
mov x0,#1
|
||||
.Lno_key:
|
||||
ret
|
||||
.size poly1305_init,.-poly1305_init
|
||||
|
||||
.type poly1305_blocks,%function
|
||||
.align 5
|
||||
poly1305_blocks:
|
||||
.Lpoly1305_blocks:
|
||||
ands $len,$len,#-16
|
||||
b.eq .Lno_data
|
||||
|
||||
ldp $h0,$h1,[$ctx] // load hash value
|
||||
ldp $h2,x17,[$ctx,#16] // [along with is_base2_26]
|
||||
ldp $r0,$r1,[$ctx,#32] // load key value
|
||||
|
||||
#ifdef __AARCH64EB__
|
||||
lsr $d0,$h0,#32
|
||||
mov w#$d1,w#$h0
|
||||
lsr $d2,$h1,#32
|
||||
mov w15,w#$h1
|
||||
lsr x16,$h2,#32
|
||||
#else
|
||||
mov w#$d0,w#$h0
|
||||
lsr $d1,$h0,#32
|
||||
mov w#$d2,w#$h1
|
||||
lsr x15,$h1,#32
|
||||
mov w16,w#$h2
|
||||
#endif
|
||||
|
||||
add $d0,$d0,$d1,lsl#26 // base 2^26 -> base 2^64
|
||||
lsr $d1,$d2,#12
|
||||
adds $d0,$d0,$d2,lsl#52
|
||||
add $d1,$d1,x15,lsl#14
|
||||
adc $d1,$d1,xzr
|
||||
lsr $d2,x16,#24
|
||||
adds $d1,$d1,x16,lsl#40
|
||||
adc $d2,$d2,xzr
|
||||
|
||||
cmp x17,#0 // is_base2_26?
|
||||
add $s1,$r1,$r1,lsr#2 // s1 = r1 + (r1 >> 2)
|
||||
csel $h0,$h0,$d0,eq // choose between radixes
|
||||
csel $h1,$h1,$d1,eq
|
||||
csel $h2,$h2,$d2,eq
|
||||
|
||||
.Loop:
|
||||
ldp $t0,$t1,[$inp],#16 // load input
|
||||
sub $len,$len,#16
|
||||
#ifdef __AARCH64EB__
|
||||
rev $t0,$t0
|
||||
rev $t1,$t1
|
||||
#endif
|
||||
adds $h0,$h0,$t0 // accumulate input
|
||||
adcs $h1,$h1,$t1
|
||||
|
||||
mul $d0,$h0,$r0 // h0*r0
|
||||
adc $h2,$h2,$padbit
|
||||
umulh $d1,$h0,$r0
|
||||
|
||||
mul $t0,$h1,$s1 // h1*5*r1
|
||||
umulh $t1,$h1,$s1
|
||||
|
||||
adds $d0,$d0,$t0
|
||||
mul $t0,$h0,$r1 // h0*r1
|
||||
adc $d1,$d1,$t1
|
||||
umulh $d2,$h0,$r1
|
||||
|
||||
adds $d1,$d1,$t0
|
||||
mul $t0,$h1,$r0 // h1*r0
|
||||
adc $d2,$d2,xzr
|
||||
umulh $t1,$h1,$r0
|
||||
|
||||
adds $d1,$d1,$t0
|
||||
mul $t0,$h2,$s1 // h2*5*r1
|
||||
adc $d2,$d2,$t1
|
||||
mul $t1,$h2,$r0 // h2*r0
|
||||
|
||||
adds $d1,$d1,$t0
|
||||
adc $d2,$d2,$t1
|
||||
|
||||
and $t0,$d2,#-4 // final reduction
|
||||
and $h2,$d2,#3
|
||||
add $t0,$t0,$d2,lsr#2
|
||||
adds $h0,$d0,$t0
|
||||
adcs $h1,$d1,xzr
|
||||
adc $h2,$h2,xzr
|
||||
|
||||
cbnz $len,.Loop
|
||||
|
||||
stp $h0,$h1,[$ctx] // store hash value
|
||||
stp $h2,xzr,[$ctx,#16] // [and clear is_base2_26]
|
||||
|
||||
.Lno_data:
|
||||
ret
|
||||
.size poly1305_blocks,.-poly1305_blocks
|
||||
|
||||
.type poly1305_emit,%function
|
||||
.align 5
|
||||
poly1305_emit:
|
||||
.Lpoly1305_emit:
|
||||
ldp $h0,$h1,[$ctx] // load hash base 2^64
|
||||
ldp $h2,$r0,[$ctx,#16] // [along with is_base2_26]
|
||||
ldp $t0,$t1,[$nonce] // load nonce
|
||||
|
||||
#ifdef __AARCH64EB__
|
||||
lsr $d0,$h0,#32
|
||||
mov w#$d1,w#$h0
|
||||
lsr $d2,$h1,#32
|
||||
mov w15,w#$h1
|
||||
lsr x16,$h2,#32
|
||||
#else
|
||||
mov w#$d0,w#$h0
|
||||
lsr $d1,$h0,#32
|
||||
mov w#$d2,w#$h1
|
||||
lsr x15,$h1,#32
|
||||
mov w16,w#$h2
|
||||
#endif
|
||||
|
||||
add $d0,$d0,$d1,lsl#26 // base 2^26 -> base 2^64
|
||||
lsr $d1,$d2,#12
|
||||
adds $d0,$d0,$d2,lsl#52
|
||||
add $d1,$d1,x15,lsl#14
|
||||
adc $d1,$d1,xzr
|
||||
lsr $d2,x16,#24
|
||||
adds $d1,$d1,x16,lsl#40
|
||||
adc $d2,$d2,xzr
|
||||
|
||||
cmp $r0,#0 // is_base2_26?
|
||||
csel $h0,$h0,$d0,eq // choose between radixes
|
||||
csel $h1,$h1,$d1,eq
|
||||
csel $h2,$h2,$d2,eq
|
||||
|
||||
adds $d0,$h0,#5 // compare to modulus
|
||||
adcs $d1,$h1,xzr
|
||||
adc $d2,$h2,xzr
|
||||
|
||||
tst $d2,#-4 // see if it's carried/borrowed
|
||||
|
||||
csel $h0,$h0,$d0,eq
|
||||
csel $h1,$h1,$d1,eq
|
||||
|
||||
#ifdef __AARCH64EB__
|
||||
ror $t0,$t0,#32 // flip nonce words
|
||||
ror $t1,$t1,#32
|
||||
#endif
|
||||
adds $h0,$h0,$t0 // accumulate nonce
|
||||
adc $h1,$h1,$t1
|
||||
#ifdef __AARCH64EB__
|
||||
rev $h0,$h0 // flip output bytes
|
||||
rev $h1,$h1
|
||||
#endif
|
||||
stp $h0,$h1,[$mac] // write result
|
||||
|
||||
ret
|
||||
.size poly1305_emit,.-poly1305_emit
|
||||
___
|
||||
my ($R0,$R1,$S1,$R2,$S2,$R3,$S3,$R4,$S4) = map("v$_.4s",(0..8));
|
||||
my ($IN01_0,$IN01_1,$IN01_2,$IN01_3,$IN01_4) = map("v$_.2s",(9..13));
|
||||
my ($IN23_0,$IN23_1,$IN23_2,$IN23_3,$IN23_4) = map("v$_.2s",(14..18));
|
||||
my ($ACC0,$ACC1,$ACC2,$ACC3,$ACC4) = map("v$_.2d",(19..23));
|
||||
my ($H0,$H1,$H2,$H3,$H4) = map("v$_.2s",(24..28));
|
||||
my ($T0,$T1,$MASK) = map("v$_",(29..31));
|
||||
|
||||
my ($in2,$zeros)=("x16","x17");
|
||||
my $is_base2_26 = $zeros; # borrow
|
||||
|
||||
$code.=<<___;
|
||||
.type poly1305_mult,%function
|
||||
.align 5
|
||||
poly1305_mult:
|
||||
mul $d0,$h0,$r0 // h0*r0
|
||||
umulh $d1,$h0,$r0
|
||||
|
||||
mul $t0,$h1,$s1 // h1*5*r1
|
||||
umulh $t1,$h1,$s1
|
||||
|
||||
adds $d0,$d0,$t0
|
||||
mul $t0,$h0,$r1 // h0*r1
|
||||
adc $d1,$d1,$t1
|
||||
umulh $d2,$h0,$r1
|
||||
|
||||
adds $d1,$d1,$t0
|
||||
mul $t0,$h1,$r0 // h1*r0
|
||||
adc $d2,$d2,xzr
|
||||
umulh $t1,$h1,$r0
|
||||
|
||||
adds $d1,$d1,$t0
|
||||
mul $t0,$h2,$s1 // h2*5*r1
|
||||
adc $d2,$d2,$t1
|
||||
mul $t1,$h2,$r0 // h2*r0
|
||||
|
||||
adds $d1,$d1,$t0
|
||||
adc $d2,$d2,$t1
|
||||
|
||||
and $t0,$d2,#-4 // final reduction
|
||||
and $h2,$d2,#3
|
||||
add $t0,$t0,$d2,lsr#2
|
||||
adds $h0,$d0,$t0
|
||||
adcs $h1,$d1,xzr
|
||||
adc $h2,$h2,xzr
|
||||
|
||||
ret
|
||||
.size poly1305_mult,.-poly1305_mult
|
||||
|
||||
.type poly1305_splat,%function
|
||||
.align 4
|
||||
poly1305_splat:
|
||||
and x12,$h0,#0x03ffffff // base 2^64 -> base 2^26
|
||||
ubfx x13,$h0,#26,#26
|
||||
extr x14,$h1,$h0,#52
|
||||
and x14,x14,#0x03ffffff
|
||||
ubfx x15,$h1,#14,#26
|
||||
extr x16,$h2,$h1,#40
|
||||
|
||||
str w12,[$ctx,#16*0] // r0
|
||||
add w12,w13,w13,lsl#2 // r1*5
|
||||
str w13,[$ctx,#16*1] // r1
|
||||
add w13,w14,w14,lsl#2 // r2*5
|
||||
str w12,[$ctx,#16*2] // s1
|
||||
str w14,[$ctx,#16*3] // r2
|
||||
add w14,w15,w15,lsl#2 // r3*5
|
||||
str w13,[$ctx,#16*4] // s2
|
||||
str w15,[$ctx,#16*5] // r3
|
||||
add w15,w16,w16,lsl#2 // r4*5
|
||||
str w14,[$ctx,#16*6] // s3
|
||||
str w16,[$ctx,#16*7] // r4
|
||||
str w15,[$ctx,#16*8] // s4
|
||||
|
||||
ret
|
||||
.size poly1305_splat,.-poly1305_splat
|
||||
|
||||
#ifdef __KERNEL__
|
||||
.globl poly1305_blocks_neon
|
||||
#endif
|
||||
.type poly1305_blocks_neon,%function
|
||||
.align 5
|
||||
poly1305_blocks_neon:
|
||||
.Lpoly1305_blocks_neon:
|
||||
ldr $is_base2_26,[$ctx,#24]
|
||||
cmp $len,#128
|
||||
b.lo .Lpoly1305_blocks
|
||||
|
||||
.inst 0xd503233f // paciasp
|
||||
stp x29,x30,[sp,#-80]!
|
||||
add x29,sp,#0
|
||||
|
||||
stp d8,d9,[sp,#16] // meet ABI requirements
|
||||
stp d10,d11,[sp,#32]
|
||||
stp d12,d13,[sp,#48]
|
||||
stp d14,d15,[sp,#64]
|
||||
|
||||
cbz $is_base2_26,.Lbase2_64_neon
|
||||
|
||||
ldp w10,w11,[$ctx] // load hash value base 2^26
|
||||
ldp w12,w13,[$ctx,#8]
|
||||
ldr w14,[$ctx,#16]
|
||||
|
||||
tst $len,#31
|
||||
b.eq .Leven_neon
|
||||
|
||||
ldp $r0,$r1,[$ctx,#32] // load key value
|
||||
|
||||
add $h0,x10,x11,lsl#26 // base 2^26 -> base 2^64
|
||||
lsr $h1,x12,#12
|
||||
adds $h0,$h0,x12,lsl#52
|
||||
add $h1,$h1,x13,lsl#14
|
||||
adc $h1,$h1,xzr
|
||||
lsr $h2,x14,#24
|
||||
adds $h1,$h1,x14,lsl#40
|
||||
adc $d2,$h2,xzr // can be partially reduced...
|
||||
|
||||
ldp $d0,$d1,[$inp],#16 // load input
|
||||
sub $len,$len,#16
|
||||
add $s1,$r1,$r1,lsr#2 // s1 = r1 + (r1 >> 2)
|
||||
|
||||
#ifdef __AARCH64EB__
|
||||
rev $d0,$d0
|
||||
rev $d1,$d1
|
||||
#endif
|
||||
adds $h0,$h0,$d0 // accumulate input
|
||||
adcs $h1,$h1,$d1
|
||||
adc $h2,$h2,$padbit
|
||||
|
||||
bl poly1305_mult
|
||||
|
||||
and x10,$h0,#0x03ffffff // base 2^64 -> base 2^26
|
||||
ubfx x11,$h0,#26,#26
|
||||
extr x12,$h1,$h0,#52
|
||||
and x12,x12,#0x03ffffff
|
||||
ubfx x13,$h1,#14,#26
|
||||
extr x14,$h2,$h1,#40
|
||||
|
||||
b .Leven_neon
|
||||
|
||||
.align 4
|
||||
.Lbase2_64_neon:
|
||||
ldp $r0,$r1,[$ctx,#32] // load key value
|
||||
|
||||
ldp $h0,$h1,[$ctx] // load hash value base 2^64
|
||||
ldr $h2,[$ctx,#16]
|
||||
|
||||
tst $len,#31
|
||||
b.eq .Linit_neon
|
||||
|
||||
ldp $d0,$d1,[$inp],#16 // load input
|
||||
sub $len,$len,#16
|
||||
add $s1,$r1,$r1,lsr#2 // s1 = r1 + (r1 >> 2)
|
||||
#ifdef __AARCH64EB__
|
||||
rev $d0,$d0
|
||||
rev $d1,$d1
|
||||
#endif
|
||||
adds $h0,$h0,$d0 // accumulate input
|
||||
adcs $h1,$h1,$d1
|
||||
adc $h2,$h2,$padbit
|
||||
|
||||
bl poly1305_mult
|
||||
|
||||
.Linit_neon:
|
||||
ldr w17,[$ctx,#48] // first table element
|
||||
and x10,$h0,#0x03ffffff // base 2^64 -> base 2^26
|
||||
ubfx x11,$h0,#26,#26
|
||||
extr x12,$h1,$h0,#52
|
||||
and x12,x12,#0x03ffffff
|
||||
ubfx x13,$h1,#14,#26
|
||||
extr x14,$h2,$h1,#40
|
||||
|
||||
cmp w17,#-1 // is value impossible?
|
||||
b.ne .Leven_neon
|
||||
|
||||
fmov ${H0},x10
|
||||
fmov ${H1},x11
|
||||
fmov ${H2},x12
|
||||
fmov ${H3},x13
|
||||
fmov ${H4},x14
|
||||
|
||||
////////////////////////////////// initialize r^n table
|
||||
mov $h0,$r0 // r^1
|
||||
add $s1,$r1,$r1,lsr#2 // s1 = r1 + (r1 >> 2)
|
||||
mov $h1,$r1
|
||||
mov $h2,xzr
|
||||
add $ctx,$ctx,#48+12
|
||||
bl poly1305_splat
|
||||
|
||||
bl poly1305_mult // r^2
|
||||
sub $ctx,$ctx,#4
|
||||
bl poly1305_splat
|
||||
|
||||
bl poly1305_mult // r^3
|
||||
sub $ctx,$ctx,#4
|
||||
bl poly1305_splat
|
||||
|
||||
bl poly1305_mult // r^4
|
||||
sub $ctx,$ctx,#4
|
||||
bl poly1305_splat
|
||||
sub $ctx,$ctx,#48 // restore original $ctx
|
||||
b .Ldo_neon
|
||||
|
||||
.align 4
|
||||
.Leven_neon:
|
||||
fmov ${H0},x10
|
||||
fmov ${H1},x11
|
||||
fmov ${H2},x12
|
||||
fmov ${H3},x13
|
||||
fmov ${H4},x14
|
||||
|
||||
.Ldo_neon:
|
||||
ldp x8,x12,[$inp,#32] // inp[2:3]
|
||||
subs $len,$len,#64
|
||||
ldp x9,x13,[$inp,#48]
|
||||
add $in2,$inp,#96
|
||||
adr $zeros,.Lzeros
|
||||
|
||||
lsl $padbit,$padbit,#24
|
||||
add x15,$ctx,#48
|
||||
|
||||
#ifdef __AARCH64EB__
|
||||
rev x8,x8
|
||||
rev x12,x12
|
||||
rev x9,x9
|
||||
rev x13,x13
|
||||
#endif
|
||||
and x4,x8,#0x03ffffff // base 2^64 -> base 2^26
|
||||
and x5,x9,#0x03ffffff
|
||||
ubfx x6,x8,#26,#26
|
||||
ubfx x7,x9,#26,#26
|
||||
add x4,x4,x5,lsl#32 // bfi x4,x5,#32,#32
|
||||
extr x8,x12,x8,#52
|
||||
extr x9,x13,x9,#52
|
||||
add x6,x6,x7,lsl#32 // bfi x6,x7,#32,#32
|
||||
fmov $IN23_0,x4
|
||||
and x8,x8,#0x03ffffff
|
||||
and x9,x9,#0x03ffffff
|
||||
ubfx x10,x12,#14,#26
|
||||
ubfx x11,x13,#14,#26
|
||||
add x12,$padbit,x12,lsr#40
|
||||
add x13,$padbit,x13,lsr#40
|
||||
add x8,x8,x9,lsl#32 // bfi x8,x9,#32,#32
|
||||
fmov $IN23_1,x6
|
||||
add x10,x10,x11,lsl#32 // bfi x10,x11,#32,#32
|
||||
add x12,x12,x13,lsl#32 // bfi x12,x13,#32,#32
|
||||
fmov $IN23_2,x8
|
||||
fmov $IN23_3,x10
|
||||
fmov $IN23_4,x12
|
||||
|
||||
ldp x8,x12,[$inp],#16 // inp[0:1]
|
||||
ldp x9,x13,[$inp],#48
|
||||
|
||||
ld1 {$R0,$R1,$S1,$R2},[x15],#64
|
||||
ld1 {$S2,$R3,$S3,$R4},[x15],#64
|
||||
ld1 {$S4},[x15]
|
||||
|
||||
#ifdef __AARCH64EB__
|
||||
rev x8,x8
|
||||
rev x12,x12
|
||||
rev x9,x9
|
||||
rev x13,x13
|
||||
#endif
|
||||
and x4,x8,#0x03ffffff // base 2^64 -> base 2^26
|
||||
and x5,x9,#0x03ffffff
|
||||
ubfx x6,x8,#26,#26
|
||||
ubfx x7,x9,#26,#26
|
||||
add x4,x4,x5,lsl#32 // bfi x4,x5,#32,#32
|
||||
extr x8,x12,x8,#52
|
||||
extr x9,x13,x9,#52
|
||||
add x6,x6,x7,lsl#32 // bfi x6,x7,#32,#32
|
||||
fmov $IN01_0,x4
|
||||
and x8,x8,#0x03ffffff
|
||||
and x9,x9,#0x03ffffff
|
||||
ubfx x10,x12,#14,#26
|
||||
ubfx x11,x13,#14,#26
|
||||
add x12,$padbit,x12,lsr#40
|
||||
add x13,$padbit,x13,lsr#40
|
||||
add x8,x8,x9,lsl#32 // bfi x8,x9,#32,#32
|
||||
fmov $IN01_1,x6
|
||||
add x10,x10,x11,lsl#32 // bfi x10,x11,#32,#32
|
||||
add x12,x12,x13,lsl#32 // bfi x12,x13,#32,#32
|
||||
movi $MASK.2d,#-1
|
||||
fmov $IN01_2,x8
|
||||
fmov $IN01_3,x10
|
||||
fmov $IN01_4,x12
|
||||
ushr $MASK.2d,$MASK.2d,#38
|
||||
|
||||
b.ls .Lskip_loop
|
||||
|
||||
.align 4
|
||||
.Loop_neon:
|
||||
////////////////////////////////////////////////////////////////
|
||||
// ((inp[0]*r^4+inp[2]*r^2+inp[4])*r^4+inp[6]*r^2
|
||||
// ((inp[1]*r^4+inp[3]*r^2+inp[5])*r^3+inp[7]*r
|
||||
// \___________________/
|
||||
// ((inp[0]*r^4+inp[2]*r^2+inp[4])*r^4+inp[6]*r^2+inp[8])*r^2
|
||||
// ((inp[1]*r^4+inp[3]*r^2+inp[5])*r^4+inp[7]*r^2+inp[9])*r
|
||||
// \___________________/ \____________________/
|
||||
//
|
||||
// Note that we start with inp[2:3]*r^2. This is because it
|
||||
// doesn't depend on reduction in previous iteration.
|
||||
////////////////////////////////////////////////////////////////
|
||||
// d4 = h0*r4 + h1*r3 + h2*r2 + h3*r1 + h4*r0
|
||||
// d3 = h0*r3 + h1*r2 + h2*r1 + h3*r0 + h4*5*r4
|
||||
// d2 = h0*r2 + h1*r1 + h2*r0 + h3*5*r4 + h4*5*r3
|
||||
// d1 = h0*r1 + h1*r0 + h2*5*r4 + h3*5*r3 + h4*5*r2
|
||||
// d0 = h0*r0 + h1*5*r4 + h2*5*r3 + h3*5*r2 + h4*5*r1
|
||||
|
||||
subs $len,$len,#64
|
||||
umull $ACC4,$IN23_0,${R4}[2]
|
||||
csel $in2,$zeros,$in2,lo
|
||||
umull $ACC3,$IN23_0,${R3}[2]
|
||||
umull $ACC2,$IN23_0,${R2}[2]
|
||||
ldp x8,x12,[$in2],#16 // inp[2:3] (or zero)
|
||||
umull $ACC1,$IN23_0,${R1}[2]
|
||||
ldp x9,x13,[$in2],#48
|
||||
umull $ACC0,$IN23_0,${R0}[2]
|
||||
#ifdef __AARCH64EB__
|
||||
rev x8,x8
|
||||
rev x12,x12
|
||||
rev x9,x9
|
||||
rev x13,x13
|
||||
#endif
|
||||
|
||||
umlal $ACC4,$IN23_1,${R3}[2]
|
||||
and x4,x8,#0x03ffffff // base 2^64 -> base 2^26
|
||||
umlal $ACC3,$IN23_1,${R2}[2]
|
||||
and x5,x9,#0x03ffffff
|
||||
umlal $ACC2,$IN23_1,${R1}[2]
|
||||
ubfx x6,x8,#26,#26
|
||||
umlal $ACC1,$IN23_1,${R0}[2]
|
||||
ubfx x7,x9,#26,#26
|
||||
umlal $ACC0,$IN23_1,${S4}[2]
|
||||
add x4,x4,x5,lsl#32 // bfi x4,x5,#32,#32
|
||||
|
||||
umlal $ACC4,$IN23_2,${R2}[2]
|
||||
extr x8,x12,x8,#52
|
||||
umlal $ACC3,$IN23_2,${R1}[2]
|
||||
extr x9,x13,x9,#52
|
||||
umlal $ACC2,$IN23_2,${R0}[2]
|
||||
add x6,x6,x7,lsl#32 // bfi x6,x7,#32,#32
|
||||
umlal $ACC1,$IN23_2,${S4}[2]
|
||||
fmov $IN23_0,x4
|
||||
umlal $ACC0,$IN23_2,${S3}[2]
|
||||
and x8,x8,#0x03ffffff
|
||||
|
||||
umlal $ACC4,$IN23_3,${R1}[2]
|
||||
and x9,x9,#0x03ffffff
|
||||
umlal $ACC3,$IN23_3,${R0}[2]
|
||||
ubfx x10,x12,#14,#26
|
||||
umlal $ACC2,$IN23_3,${S4}[2]
|
||||
ubfx x11,x13,#14,#26
|
||||
umlal $ACC1,$IN23_3,${S3}[2]
|
||||
add x8,x8,x9,lsl#32 // bfi x8,x9,#32,#32
|
||||
umlal $ACC0,$IN23_3,${S2}[2]
|
||||
fmov $IN23_1,x6
|
||||
|
||||
add $IN01_2,$IN01_2,$H2
|
||||
add x12,$padbit,x12,lsr#40
|
||||
umlal $ACC4,$IN23_4,${R0}[2]
|
||||
add x13,$padbit,x13,lsr#40
|
||||
umlal $ACC3,$IN23_4,${S4}[2]
|
||||
add x10,x10,x11,lsl#32 // bfi x10,x11,#32,#32
|
||||
umlal $ACC2,$IN23_4,${S3}[2]
|
||||
add x12,x12,x13,lsl#32 // bfi x12,x13,#32,#32
|
||||
umlal $ACC1,$IN23_4,${S2}[2]
|
||||
fmov $IN23_2,x8
|
||||
umlal $ACC0,$IN23_4,${S1}[2]
|
||||
fmov $IN23_3,x10
|
||||
|
||||
////////////////////////////////////////////////////////////////
|
||||
// (hash+inp[0:1])*r^4 and accumulate
|
||||
|
||||
add $IN01_0,$IN01_0,$H0
|
||||
fmov $IN23_4,x12
|
||||
umlal $ACC3,$IN01_2,${R1}[0]
|
||||
ldp x8,x12,[$inp],#16 // inp[0:1]
|
||||
umlal $ACC0,$IN01_2,${S3}[0]
|
||||
ldp x9,x13,[$inp],#48
|
||||
umlal $ACC4,$IN01_2,${R2}[0]
|
||||
umlal $ACC1,$IN01_2,${S4}[0]
|
||||
umlal $ACC2,$IN01_2,${R0}[0]
|
||||
#ifdef __AARCH64EB__
|
||||
rev x8,x8
|
||||
rev x12,x12
|
||||
rev x9,x9
|
||||
rev x13,x13
|
||||
#endif
|
||||
|
||||
add $IN01_1,$IN01_1,$H1
|
||||
umlal $ACC3,$IN01_0,${R3}[0]
|
||||
umlal $ACC4,$IN01_0,${R4}[0]
|
||||
and x4,x8,#0x03ffffff // base 2^64 -> base 2^26
|
||||
umlal $ACC2,$IN01_0,${R2}[0]
|
||||
and x5,x9,#0x03ffffff
|
||||
umlal $ACC0,$IN01_0,${R0}[0]
|
||||
ubfx x6,x8,#26,#26
|
||||
umlal $ACC1,$IN01_0,${R1}[0]
|
||||
ubfx x7,x9,#26,#26
|
||||
|
||||
add $IN01_3,$IN01_3,$H3
|
||||
add x4,x4,x5,lsl#32 // bfi x4,x5,#32,#32
|
||||
umlal $ACC3,$IN01_1,${R2}[0]
|
||||
extr x8,x12,x8,#52
|
||||
umlal $ACC4,$IN01_1,${R3}[0]
|
||||
extr x9,x13,x9,#52
|
||||
umlal $ACC0,$IN01_1,${S4}[0]
|
||||
add x6,x6,x7,lsl#32 // bfi x6,x7,#32,#32
|
||||
umlal $ACC2,$IN01_1,${R1}[0]
|
||||
fmov $IN01_0,x4
|
||||
umlal $ACC1,$IN01_1,${R0}[0]
|
||||
and x8,x8,#0x03ffffff
|
||||
|
||||
add $IN01_4,$IN01_4,$H4
|
||||
and x9,x9,#0x03ffffff
|
||||
umlal $ACC3,$IN01_3,${R0}[0]
|
||||
ubfx x10,x12,#14,#26
|
||||
umlal $ACC0,$IN01_3,${S2}[0]
|
||||
ubfx x11,x13,#14,#26
|
||||
umlal $ACC4,$IN01_3,${R1}[0]
|
||||
add x8,x8,x9,lsl#32 // bfi x8,x9,#32,#32
|
||||
umlal $ACC1,$IN01_3,${S3}[0]
|
||||
fmov $IN01_1,x6
|
||||
umlal $ACC2,$IN01_3,${S4}[0]
|
||||
add x12,$padbit,x12,lsr#40
|
||||
|
||||
umlal $ACC3,$IN01_4,${S4}[0]
|
||||
add x13,$padbit,x13,lsr#40
|
||||
umlal $ACC0,$IN01_4,${S1}[0]
|
||||
add x10,x10,x11,lsl#32 // bfi x10,x11,#32,#32
|
||||
umlal $ACC4,$IN01_4,${R0}[0]
|
||||
add x12,x12,x13,lsl#32 // bfi x12,x13,#32,#32
|
||||
umlal $ACC1,$IN01_4,${S2}[0]
|
||||
fmov $IN01_2,x8
|
||||
umlal $ACC2,$IN01_4,${S3}[0]
|
||||
fmov $IN01_3,x10
|
||||
fmov $IN01_4,x12
|
||||
|
||||
/////////////////////////////////////////////////////////////////
|
||||
// lazy reduction as discussed in "NEON crypto" by D.J. Bernstein
|
||||
// and P. Schwabe
|
||||
//
|
||||
// [see discussion in poly1305-armv4 module]
|
||||
|
||||
ushr $T0.2d,$ACC3,#26
|
||||
xtn $H3,$ACC3
|
||||
ushr $T1.2d,$ACC0,#26
|
||||
and $ACC0,$ACC0,$MASK.2d
|
||||
add $ACC4,$ACC4,$T0.2d // h3 -> h4
|
||||
bic $H3,#0xfc,lsl#24 // &=0x03ffffff
|
||||
add $ACC1,$ACC1,$T1.2d // h0 -> h1
|
||||
|
||||
ushr $T0.2d,$ACC4,#26
|
||||
xtn $H4,$ACC4
|
||||
ushr $T1.2d,$ACC1,#26
|
||||
xtn $H1,$ACC1
|
||||
bic $H4,#0xfc,lsl#24
|
||||
add $ACC2,$ACC2,$T1.2d // h1 -> h2
|
||||
|
||||
add $ACC0,$ACC0,$T0.2d
|
||||
shl $T0.2d,$T0.2d,#2
|
||||
shrn $T1.2s,$ACC2,#26
|
||||
xtn $H2,$ACC2
|
||||
add $ACC0,$ACC0,$T0.2d // h4 -> h0
|
||||
bic $H1,#0xfc,lsl#24
|
||||
add $H3,$H3,$T1.2s // h2 -> h3
|
||||
bic $H2,#0xfc,lsl#24
|
||||
|
||||
shrn $T0.2s,$ACC0,#26
|
||||
xtn $H0,$ACC0
|
||||
ushr $T1.2s,$H3,#26
|
||||
bic $H3,#0xfc,lsl#24
|
||||
bic $H0,#0xfc,lsl#24
|
||||
add $H1,$H1,$T0.2s // h0 -> h1
|
||||
add $H4,$H4,$T1.2s // h3 -> h4
|
||||
|
||||
b.hi .Loop_neon
|
||||
|
||||
.Lskip_loop:
|
||||
dup $IN23_2,${IN23_2}[0]
|
||||
add $IN01_2,$IN01_2,$H2
|
||||
|
||||
////////////////////////////////////////////////////////////////
|
||||
// multiply (inp[0:1]+hash) or inp[2:3] by r^2:r^1
|
||||
|
||||
adds $len,$len,#32
|
||||
b.ne .Long_tail
|
||||
|
||||
dup $IN23_2,${IN01_2}[0]
|
||||
add $IN23_0,$IN01_0,$H0
|
||||
add $IN23_3,$IN01_3,$H3
|
||||
add $IN23_1,$IN01_1,$H1
|
||||
add $IN23_4,$IN01_4,$H4
|
||||
|
||||
.Long_tail:
|
||||
dup $IN23_0,${IN23_0}[0]
|
||||
umull2 $ACC0,$IN23_2,${S3}
|
||||
umull2 $ACC3,$IN23_2,${R1}
|
||||
umull2 $ACC4,$IN23_2,${R2}
|
||||
umull2 $ACC2,$IN23_2,${R0}
|
||||
umull2 $ACC1,$IN23_2,${S4}
|
||||
|
||||
dup $IN23_1,${IN23_1}[0]
|
||||
umlal2 $ACC0,$IN23_0,${R0}
|
||||
umlal2 $ACC2,$IN23_0,${R2}
|
||||
umlal2 $ACC3,$IN23_0,${R3}
|
||||
umlal2 $ACC4,$IN23_0,${R4}
|
||||
umlal2 $ACC1,$IN23_0,${R1}
|
||||
|
||||
dup $IN23_3,${IN23_3}[0]
|
||||
umlal2 $ACC0,$IN23_1,${S4}
|
||||
umlal2 $ACC3,$IN23_1,${R2}
|
||||
umlal2 $ACC2,$IN23_1,${R1}
|
||||
umlal2 $ACC4,$IN23_1,${R3}
|
||||
umlal2 $ACC1,$IN23_1,${R0}
|
||||
|
||||
dup $IN23_4,${IN23_4}[0]
|
||||
umlal2 $ACC3,$IN23_3,${R0}
|
||||
umlal2 $ACC4,$IN23_3,${R1}
|
||||
umlal2 $ACC0,$IN23_3,${S2}
|
||||
umlal2 $ACC1,$IN23_3,${S3}
|
||||
umlal2 $ACC2,$IN23_3,${S4}
|
||||
|
||||
umlal2 $ACC3,$IN23_4,${S4}
|
||||
umlal2 $ACC0,$IN23_4,${S1}
|
||||
umlal2 $ACC4,$IN23_4,${R0}
|
||||
umlal2 $ACC1,$IN23_4,${S2}
|
||||
umlal2 $ACC2,$IN23_4,${S3}
|
||||
|
||||
b.eq .Lshort_tail
|
||||
|
||||
////////////////////////////////////////////////////////////////
|
||||
// (hash+inp[0:1])*r^4:r^3 and accumulate
|
||||
|
||||
add $IN01_0,$IN01_0,$H0
|
||||
umlal $ACC3,$IN01_2,${R1}
|
||||
umlal $ACC0,$IN01_2,${S3}
|
||||
umlal $ACC4,$IN01_2,${R2}
|
||||
umlal $ACC1,$IN01_2,${S4}
|
||||
umlal $ACC2,$IN01_2,${R0}
|
||||
|
||||
add $IN01_1,$IN01_1,$H1
|
||||
umlal $ACC3,$IN01_0,${R3}
|
||||
umlal $ACC0,$IN01_0,${R0}
|
||||
umlal $ACC4,$IN01_0,${R4}
|
||||
umlal $ACC1,$IN01_0,${R1}
|
||||
umlal $ACC2,$IN01_0,${R2}
|
||||
|
||||
add $IN01_3,$IN01_3,$H3
|
||||
umlal $ACC3,$IN01_1,${R2}
|
||||
umlal $ACC0,$IN01_1,${S4}
|
||||
umlal $ACC4,$IN01_1,${R3}
|
||||
umlal $ACC1,$IN01_1,${R0}
|
||||
umlal $ACC2,$IN01_1,${R1}
|
||||
|
||||
add $IN01_4,$IN01_4,$H4
|
||||
umlal $ACC3,$IN01_3,${R0}
|
||||
umlal $ACC0,$IN01_3,${S2}
|
||||
umlal $ACC4,$IN01_3,${R1}
|
||||
umlal $ACC1,$IN01_3,${S3}
|
||||
umlal $ACC2,$IN01_3,${S4}
|
||||
|
||||
umlal $ACC3,$IN01_4,${S4}
|
||||
umlal $ACC0,$IN01_4,${S1}
|
||||
umlal $ACC4,$IN01_4,${R0}
|
||||
umlal $ACC1,$IN01_4,${S2}
|
||||
umlal $ACC2,$IN01_4,${S3}
|
||||
|
||||
.Lshort_tail:
|
||||
////////////////////////////////////////////////////////////////
|
||||
// horizontal add
|
||||
|
||||
addp $ACC3,$ACC3,$ACC3
|
||||
ldp d8,d9,[sp,#16] // meet ABI requirements
|
||||
addp $ACC0,$ACC0,$ACC0
|
||||
ldp d10,d11,[sp,#32]
|
||||
addp $ACC4,$ACC4,$ACC4
|
||||
ldp d12,d13,[sp,#48]
|
||||
addp $ACC1,$ACC1,$ACC1
|
||||
ldp d14,d15,[sp,#64]
|
||||
addp $ACC2,$ACC2,$ACC2
|
||||
ldr x30,[sp,#8]
|
||||
|
||||
////////////////////////////////////////////////////////////////
|
||||
// lazy reduction, but without narrowing
|
||||
|
||||
ushr $T0.2d,$ACC3,#26
|
||||
and $ACC3,$ACC3,$MASK.2d
|
||||
ushr $T1.2d,$ACC0,#26
|
||||
and $ACC0,$ACC0,$MASK.2d
|
||||
|
||||
add $ACC4,$ACC4,$T0.2d // h3 -> h4
|
||||
add $ACC1,$ACC1,$T1.2d // h0 -> h1
|
||||
|
||||
ushr $T0.2d,$ACC4,#26
|
||||
and $ACC4,$ACC4,$MASK.2d
|
||||
ushr $T1.2d,$ACC1,#26
|
||||
and $ACC1,$ACC1,$MASK.2d
|
||||
add $ACC2,$ACC2,$T1.2d // h1 -> h2
|
||||
|
||||
add $ACC0,$ACC0,$T0.2d
|
||||
shl $T0.2d,$T0.2d,#2
|
||||
ushr $T1.2d,$ACC2,#26
|
||||
and $ACC2,$ACC2,$MASK.2d
|
||||
add $ACC0,$ACC0,$T0.2d // h4 -> h0
|
||||
add $ACC3,$ACC3,$T1.2d // h2 -> h3
|
||||
|
||||
ushr $T0.2d,$ACC0,#26
|
||||
and $ACC0,$ACC0,$MASK.2d
|
||||
ushr $T1.2d,$ACC3,#26
|
||||
and $ACC3,$ACC3,$MASK.2d
|
||||
add $ACC1,$ACC1,$T0.2d // h0 -> h1
|
||||
add $ACC4,$ACC4,$T1.2d // h3 -> h4
|
||||
|
||||
////////////////////////////////////////////////////////////////
|
||||
// write the result, can be partially reduced
|
||||
|
||||
st4 {$ACC0,$ACC1,$ACC2,$ACC3}[0],[$ctx],#16
|
||||
mov x4,#1
|
||||
st1 {$ACC4}[0],[$ctx]
|
||||
str x4,[$ctx,#8] // set is_base2_26
|
||||
|
||||
ldr x29,[sp],#80
|
||||
.inst 0xd50323bf // autiasp
|
||||
ret
|
||||
.size poly1305_blocks_neon,.-poly1305_blocks_neon
|
||||
|
||||
.align 5
|
||||
.Lzeros:
|
||||
.long 0,0,0,0,0,0,0,0
|
||||
.asciz "Poly1305 for ARMv8, CRYPTOGAMS by \@dot-asm"
|
||||
.align 2
|
||||
#if !defined(__KERNEL__) && !defined(_WIN64)
|
||||
.comm OPENSSL_armcap_P,4,4
|
||||
.hidden OPENSSL_armcap_P
|
||||
#endif
|
||||
___
|
||||
|
||||
foreach (split("\n",$code)) {
|
||||
s/\b(shrn\s+v[0-9]+)\.[24]d/$1.2s/ or
|
||||
s/\b(fmov\s+)v([0-9]+)[^,]*,\s*x([0-9]+)/$1d$2,x$3/ or
|
||||
(m/\bdup\b/ and (s/\.[24]s/.2d/g or 1)) or
|
||||
(m/\b(eor|and)/ and (s/\.[248][sdh]/.16b/g or 1)) or
|
||||
(m/\bum(ul|la)l\b/ and (s/\.4s/.2s/g or 1)) or
|
||||
(m/\bum(ul|la)l2\b/ and (s/\.2s/.4s/g or 1)) or
|
||||
(m/\bst[1-4]\s+{[^}]+}\[/ and (s/\.[24]d/.s/g or 1));
|
||||
|
||||
s/\.[124]([sd])\[/.$1\[/;
|
||||
s/w#x([0-9]+)/w$1/g;
|
||||
|
||||
print $_,"\n";
|
||||
}
|
||||
close STDOUT;
|
||||
835
arch/arm64/crypto/poly1305-core.S_shipped
Normal file
835
arch/arm64/crypto/poly1305-core.S_shipped
Normal file
|
|
@ -0,0 +1,835 @@
|
|||
#ifndef __KERNEL__
|
||||
# include "arm_arch.h"
|
||||
.extern OPENSSL_armcap_P
|
||||
#endif
|
||||
|
||||
.text
|
||||
|
||||
// forward "declarations" are required for Apple
|
||||
.globl poly1305_blocks
|
||||
.globl poly1305_emit
|
||||
|
||||
.globl poly1305_init
|
||||
.type poly1305_init,%function
|
||||
.align 5
|
||||
poly1305_init:
|
||||
cmp x1,xzr
|
||||
stp xzr,xzr,[x0] // zero hash value
|
||||
stp xzr,xzr,[x0,#16] // [along with is_base2_26]
|
||||
|
||||
csel x0,xzr,x0,eq
|
||||
b.eq .Lno_key
|
||||
|
||||
#ifndef __KERNEL__
|
||||
adrp x17,OPENSSL_armcap_P
|
||||
ldr w17,[x17,#:lo12:OPENSSL_armcap_P]
|
||||
#endif
|
||||
|
||||
ldp x7,x8,[x1] // load key
|
||||
mov x9,#0xfffffffc0fffffff
|
||||
movk x9,#0x0fff,lsl#48
|
||||
#ifdef __AARCH64EB__
|
||||
rev x7,x7 // flip bytes
|
||||
rev x8,x8
|
||||
#endif
|
||||
and x7,x7,x9 // &=0ffffffc0fffffff
|
||||
and x9,x9,#-4
|
||||
and x8,x8,x9 // &=0ffffffc0ffffffc
|
||||
mov w9,#-1
|
||||
stp x7,x8,[x0,#32] // save key value
|
||||
str w9,[x0,#48] // impossible key power value
|
||||
|
||||
#ifndef __KERNEL__
|
||||
tst w17,#ARMV7_NEON
|
||||
|
||||
adr x12,.Lpoly1305_blocks
|
||||
adr x7,.Lpoly1305_blocks_neon
|
||||
adr x13,.Lpoly1305_emit
|
||||
|
||||
csel x12,x12,x7,eq
|
||||
|
||||
# ifdef __ILP32__
|
||||
stp w12,w13,[x2]
|
||||
# else
|
||||
stp x12,x13,[x2]
|
||||
# endif
|
||||
#endif
|
||||
mov x0,#1
|
||||
.Lno_key:
|
||||
ret
|
||||
.size poly1305_init,.-poly1305_init
|
||||
|
||||
.type poly1305_blocks,%function
|
||||
.align 5
|
||||
poly1305_blocks:
|
||||
.Lpoly1305_blocks:
|
||||
ands x2,x2,#-16
|
||||
b.eq .Lno_data
|
||||
|
||||
ldp x4,x5,[x0] // load hash value
|
||||
ldp x6,x17,[x0,#16] // [along with is_base2_26]
|
||||
ldp x7,x8,[x0,#32] // load key value
|
||||
|
||||
#ifdef __AARCH64EB__
|
||||
lsr x12,x4,#32
|
||||
mov w13,w4
|
||||
lsr x14,x5,#32
|
||||
mov w15,w5
|
||||
lsr x16,x6,#32
|
||||
#else
|
||||
mov w12,w4
|
||||
lsr x13,x4,#32
|
||||
mov w14,w5
|
||||
lsr x15,x5,#32
|
||||
mov w16,w6
|
||||
#endif
|
||||
|
||||
add x12,x12,x13,lsl#26 // base 2^26 -> base 2^64
|
||||
lsr x13,x14,#12
|
||||
adds x12,x12,x14,lsl#52
|
||||
add x13,x13,x15,lsl#14
|
||||
adc x13,x13,xzr
|
||||
lsr x14,x16,#24
|
||||
adds x13,x13,x16,lsl#40
|
||||
adc x14,x14,xzr
|
||||
|
||||
cmp x17,#0 // is_base2_26?
|
||||
add x9,x8,x8,lsr#2 // s1 = r1 + (r1 >> 2)
|
||||
csel x4,x4,x12,eq // choose between radixes
|
||||
csel x5,x5,x13,eq
|
||||
csel x6,x6,x14,eq
|
||||
|
||||
.Loop:
|
||||
ldp x10,x11,[x1],#16 // load input
|
||||
sub x2,x2,#16
|
||||
#ifdef __AARCH64EB__
|
||||
rev x10,x10
|
||||
rev x11,x11
|
||||
#endif
|
||||
adds x4,x4,x10 // accumulate input
|
||||
adcs x5,x5,x11
|
||||
|
||||
mul x12,x4,x7 // h0*r0
|
||||
adc x6,x6,x3
|
||||
umulh x13,x4,x7
|
||||
|
||||
mul x10,x5,x9 // h1*5*r1
|
||||
umulh x11,x5,x9
|
||||
|
||||
adds x12,x12,x10
|
||||
mul x10,x4,x8 // h0*r1
|
||||
adc x13,x13,x11
|
||||
umulh x14,x4,x8
|
||||
|
||||
adds x13,x13,x10
|
||||
mul x10,x5,x7 // h1*r0
|
||||
adc x14,x14,xzr
|
||||
umulh x11,x5,x7
|
||||
|
||||
adds x13,x13,x10
|
||||
mul x10,x6,x9 // h2*5*r1
|
||||
adc x14,x14,x11
|
||||
mul x11,x6,x7 // h2*r0
|
||||
|
||||
adds x13,x13,x10
|
||||
adc x14,x14,x11
|
||||
|
||||
and x10,x14,#-4 // final reduction
|
||||
and x6,x14,#3
|
||||
add x10,x10,x14,lsr#2
|
||||
adds x4,x12,x10
|
||||
adcs x5,x13,xzr
|
||||
adc x6,x6,xzr
|
||||
|
||||
cbnz x2,.Loop
|
||||
|
||||
stp x4,x5,[x0] // store hash value
|
||||
stp x6,xzr,[x0,#16] // [and clear is_base2_26]
|
||||
|
||||
.Lno_data:
|
||||
ret
|
||||
.size poly1305_blocks,.-poly1305_blocks
|
||||
|
||||
.type poly1305_emit,%function
|
||||
.align 5
|
||||
poly1305_emit:
|
||||
.Lpoly1305_emit:
|
||||
ldp x4,x5,[x0] // load hash base 2^64
|
||||
ldp x6,x7,[x0,#16] // [along with is_base2_26]
|
||||
ldp x10,x11,[x2] // load nonce
|
||||
|
||||
#ifdef __AARCH64EB__
|
||||
lsr x12,x4,#32
|
||||
mov w13,w4
|
||||
lsr x14,x5,#32
|
||||
mov w15,w5
|
||||
lsr x16,x6,#32
|
||||
#else
|
||||
mov w12,w4
|
||||
lsr x13,x4,#32
|
||||
mov w14,w5
|
||||
lsr x15,x5,#32
|
||||
mov w16,w6
|
||||
#endif
|
||||
|
||||
add x12,x12,x13,lsl#26 // base 2^26 -> base 2^64
|
||||
lsr x13,x14,#12
|
||||
adds x12,x12,x14,lsl#52
|
||||
add x13,x13,x15,lsl#14
|
||||
adc x13,x13,xzr
|
||||
lsr x14,x16,#24
|
||||
adds x13,x13,x16,lsl#40
|
||||
adc x14,x14,xzr
|
||||
|
||||
cmp x7,#0 // is_base2_26?
|
||||
csel x4,x4,x12,eq // choose between radixes
|
||||
csel x5,x5,x13,eq
|
||||
csel x6,x6,x14,eq
|
||||
|
||||
adds x12,x4,#5 // compare to modulus
|
||||
adcs x13,x5,xzr
|
||||
adc x14,x6,xzr
|
||||
|
||||
tst x14,#-4 // see if it's carried/borrowed
|
||||
|
||||
csel x4,x4,x12,eq
|
||||
csel x5,x5,x13,eq
|
||||
|
||||
#ifdef __AARCH64EB__
|
||||
ror x10,x10,#32 // flip nonce words
|
||||
ror x11,x11,#32
|
||||
#endif
|
||||
adds x4,x4,x10 // accumulate nonce
|
||||
adc x5,x5,x11
|
||||
#ifdef __AARCH64EB__
|
||||
rev x4,x4 // flip output bytes
|
||||
rev x5,x5
|
||||
#endif
|
||||
stp x4,x5,[x1] // write result
|
||||
|
||||
ret
|
||||
.size poly1305_emit,.-poly1305_emit
|
||||
.type poly1305_mult,%function
|
||||
.align 5
|
||||
poly1305_mult:
|
||||
mul x12,x4,x7 // h0*r0
|
||||
umulh x13,x4,x7
|
||||
|
||||
mul x10,x5,x9 // h1*5*r1
|
||||
umulh x11,x5,x9
|
||||
|
||||
adds x12,x12,x10
|
||||
mul x10,x4,x8 // h0*r1
|
||||
adc x13,x13,x11
|
||||
umulh x14,x4,x8
|
||||
|
||||
adds x13,x13,x10
|
||||
mul x10,x5,x7 // h1*r0
|
||||
adc x14,x14,xzr
|
||||
umulh x11,x5,x7
|
||||
|
||||
adds x13,x13,x10
|
||||
mul x10,x6,x9 // h2*5*r1
|
||||
adc x14,x14,x11
|
||||
mul x11,x6,x7 // h2*r0
|
||||
|
||||
adds x13,x13,x10
|
||||
adc x14,x14,x11
|
||||
|
||||
and x10,x14,#-4 // final reduction
|
||||
and x6,x14,#3
|
||||
add x10,x10,x14,lsr#2
|
||||
adds x4,x12,x10
|
||||
adcs x5,x13,xzr
|
||||
adc x6,x6,xzr
|
||||
|
||||
ret
|
||||
.size poly1305_mult,.-poly1305_mult
|
||||
|
||||
.type poly1305_splat,%function
|
||||
.align 4
|
||||
poly1305_splat:
|
||||
and x12,x4,#0x03ffffff // base 2^64 -> base 2^26
|
||||
ubfx x13,x4,#26,#26
|
||||
extr x14,x5,x4,#52
|
||||
and x14,x14,#0x03ffffff
|
||||
ubfx x15,x5,#14,#26
|
||||
extr x16,x6,x5,#40
|
||||
|
||||
str w12,[x0,#16*0] // r0
|
||||
add w12,w13,w13,lsl#2 // r1*5
|
||||
str w13,[x0,#16*1] // r1
|
||||
add w13,w14,w14,lsl#2 // r2*5
|
||||
str w12,[x0,#16*2] // s1
|
||||
str w14,[x0,#16*3] // r2
|
||||
add w14,w15,w15,lsl#2 // r3*5
|
||||
str w13,[x0,#16*4] // s2
|
||||
str w15,[x0,#16*5] // r3
|
||||
add w15,w16,w16,lsl#2 // r4*5
|
||||
str w14,[x0,#16*6] // s3
|
||||
str w16,[x0,#16*7] // r4
|
||||
str w15,[x0,#16*8] // s4
|
||||
|
||||
ret
|
||||
.size poly1305_splat,.-poly1305_splat
|
||||
|
||||
#ifdef __KERNEL__
|
||||
.globl poly1305_blocks_neon
|
||||
#endif
|
||||
.type poly1305_blocks_neon,%function
|
||||
.align 5
|
||||
poly1305_blocks_neon:
|
||||
.Lpoly1305_blocks_neon:
|
||||
ldr x17,[x0,#24]
|
||||
cmp x2,#128
|
||||
b.lo .Lpoly1305_blocks
|
||||
|
||||
.inst 0xd503233f // paciasp
|
||||
stp x29,x30,[sp,#-80]!
|
||||
add x29,sp,#0
|
||||
|
||||
stp d8,d9,[sp,#16] // meet ABI requirements
|
||||
stp d10,d11,[sp,#32]
|
||||
stp d12,d13,[sp,#48]
|
||||
stp d14,d15,[sp,#64]
|
||||
|
||||
cbz x17,.Lbase2_64_neon
|
||||
|
||||
ldp w10,w11,[x0] // load hash value base 2^26
|
||||
ldp w12,w13,[x0,#8]
|
||||
ldr w14,[x0,#16]
|
||||
|
||||
tst x2,#31
|
||||
b.eq .Leven_neon
|
||||
|
||||
ldp x7,x8,[x0,#32] // load key value
|
||||
|
||||
add x4,x10,x11,lsl#26 // base 2^26 -> base 2^64
|
||||
lsr x5,x12,#12
|
||||
adds x4,x4,x12,lsl#52
|
||||
add x5,x5,x13,lsl#14
|
||||
adc x5,x5,xzr
|
||||
lsr x6,x14,#24
|
||||
adds x5,x5,x14,lsl#40
|
||||
adc x14,x6,xzr // can be partially reduced...
|
||||
|
||||
ldp x12,x13,[x1],#16 // load input
|
||||
sub x2,x2,#16
|
||||
add x9,x8,x8,lsr#2 // s1 = r1 + (r1 >> 2)
|
||||
|
||||
#ifdef __AARCH64EB__
|
||||
rev x12,x12
|
||||
rev x13,x13
|
||||
#endif
|
||||
adds x4,x4,x12 // accumulate input
|
||||
adcs x5,x5,x13
|
||||
adc x6,x6,x3
|
||||
|
||||
bl poly1305_mult
|
||||
|
||||
and x10,x4,#0x03ffffff // base 2^64 -> base 2^26
|
||||
ubfx x11,x4,#26,#26
|
||||
extr x12,x5,x4,#52
|
||||
and x12,x12,#0x03ffffff
|
||||
ubfx x13,x5,#14,#26
|
||||
extr x14,x6,x5,#40
|
||||
|
||||
b .Leven_neon
|
||||
|
||||
.align 4
|
||||
.Lbase2_64_neon:
|
||||
ldp x7,x8,[x0,#32] // load key value
|
||||
|
||||
ldp x4,x5,[x0] // load hash value base 2^64
|
||||
ldr x6,[x0,#16]
|
||||
|
||||
tst x2,#31
|
||||
b.eq .Linit_neon
|
||||
|
||||
ldp x12,x13,[x1],#16 // load input
|
||||
sub x2,x2,#16
|
||||
add x9,x8,x8,lsr#2 // s1 = r1 + (r1 >> 2)
|
||||
#ifdef __AARCH64EB__
|
||||
rev x12,x12
|
||||
rev x13,x13
|
||||
#endif
|
||||
adds x4,x4,x12 // accumulate input
|
||||
adcs x5,x5,x13
|
||||
adc x6,x6,x3
|
||||
|
||||
bl poly1305_mult
|
||||
|
||||
.Linit_neon:
|
||||
ldr w17,[x0,#48] // first table element
|
||||
and x10,x4,#0x03ffffff // base 2^64 -> base 2^26
|
||||
ubfx x11,x4,#26,#26
|
||||
extr x12,x5,x4,#52
|
||||
and x12,x12,#0x03ffffff
|
||||
ubfx x13,x5,#14,#26
|
||||
extr x14,x6,x5,#40
|
||||
|
||||
cmp w17,#-1 // is value impossible?
|
||||
b.ne .Leven_neon
|
||||
|
||||
fmov d24,x10
|
||||
fmov d25,x11
|
||||
fmov d26,x12
|
||||
fmov d27,x13
|
||||
fmov d28,x14
|
||||
|
||||
////////////////////////////////// initialize r^n table
|
||||
mov x4,x7 // r^1
|
||||
add x9,x8,x8,lsr#2 // s1 = r1 + (r1 >> 2)
|
||||
mov x5,x8
|
||||
mov x6,xzr
|
||||
add x0,x0,#48+12
|
||||
bl poly1305_splat
|
||||
|
||||
bl poly1305_mult // r^2
|
||||
sub x0,x0,#4
|
||||
bl poly1305_splat
|
||||
|
||||
bl poly1305_mult // r^3
|
||||
sub x0,x0,#4
|
||||
bl poly1305_splat
|
||||
|
||||
bl poly1305_mult // r^4
|
||||
sub x0,x0,#4
|
||||
bl poly1305_splat
|
||||
sub x0,x0,#48 // restore original x0
|
||||
b .Ldo_neon
|
||||
|
||||
.align 4
|
||||
.Leven_neon:
|
||||
fmov d24,x10
|
||||
fmov d25,x11
|
||||
fmov d26,x12
|
||||
fmov d27,x13
|
||||
fmov d28,x14
|
||||
|
||||
.Ldo_neon:
|
||||
ldp x8,x12,[x1,#32] // inp[2:3]
|
||||
subs x2,x2,#64
|
||||
ldp x9,x13,[x1,#48]
|
||||
add x16,x1,#96
|
||||
adr x17,.Lzeros
|
||||
|
||||
lsl x3,x3,#24
|
||||
add x15,x0,#48
|
||||
|
||||
#ifdef __AARCH64EB__
|
||||
rev x8,x8
|
||||
rev x12,x12
|
||||
rev x9,x9
|
||||
rev x13,x13
|
||||
#endif
|
||||
and x4,x8,#0x03ffffff // base 2^64 -> base 2^26
|
||||
and x5,x9,#0x03ffffff
|
||||
ubfx x6,x8,#26,#26
|
||||
ubfx x7,x9,#26,#26
|
||||
add x4,x4,x5,lsl#32 // bfi x4,x5,#32,#32
|
||||
extr x8,x12,x8,#52
|
||||
extr x9,x13,x9,#52
|
||||
add x6,x6,x7,lsl#32 // bfi x6,x7,#32,#32
|
||||
fmov d14,x4
|
||||
and x8,x8,#0x03ffffff
|
||||
and x9,x9,#0x03ffffff
|
||||
ubfx x10,x12,#14,#26
|
||||
ubfx x11,x13,#14,#26
|
||||
add x12,x3,x12,lsr#40
|
||||
add x13,x3,x13,lsr#40
|
||||
add x8,x8,x9,lsl#32 // bfi x8,x9,#32,#32
|
||||
fmov d15,x6
|
||||
add x10,x10,x11,lsl#32 // bfi x10,x11,#32,#32
|
||||
add x12,x12,x13,lsl#32 // bfi x12,x13,#32,#32
|
||||
fmov d16,x8
|
||||
fmov d17,x10
|
||||
fmov d18,x12
|
||||
|
||||
ldp x8,x12,[x1],#16 // inp[0:1]
|
||||
ldp x9,x13,[x1],#48
|
||||
|
||||
ld1 {v0.4s,v1.4s,v2.4s,v3.4s},[x15],#64
|
||||
ld1 {v4.4s,v5.4s,v6.4s,v7.4s},[x15],#64
|
||||
ld1 {v8.4s},[x15]
|
||||
|
||||
#ifdef __AARCH64EB__
|
||||
rev x8,x8
|
||||
rev x12,x12
|
||||
rev x9,x9
|
||||
rev x13,x13
|
||||
#endif
|
||||
and x4,x8,#0x03ffffff // base 2^64 -> base 2^26
|
||||
and x5,x9,#0x03ffffff
|
||||
ubfx x6,x8,#26,#26
|
||||
ubfx x7,x9,#26,#26
|
||||
add x4,x4,x5,lsl#32 // bfi x4,x5,#32,#32
|
||||
extr x8,x12,x8,#52
|
||||
extr x9,x13,x9,#52
|
||||
add x6,x6,x7,lsl#32 // bfi x6,x7,#32,#32
|
||||
fmov d9,x4
|
||||
and x8,x8,#0x03ffffff
|
||||
and x9,x9,#0x03ffffff
|
||||
ubfx x10,x12,#14,#26
|
||||
ubfx x11,x13,#14,#26
|
||||
add x12,x3,x12,lsr#40
|
||||
add x13,x3,x13,lsr#40
|
||||
add x8,x8,x9,lsl#32 // bfi x8,x9,#32,#32
|
||||
fmov d10,x6
|
||||
add x10,x10,x11,lsl#32 // bfi x10,x11,#32,#32
|
||||
add x12,x12,x13,lsl#32 // bfi x12,x13,#32,#32
|
||||
movi v31.2d,#-1
|
||||
fmov d11,x8
|
||||
fmov d12,x10
|
||||
fmov d13,x12
|
||||
ushr v31.2d,v31.2d,#38
|
||||
|
||||
b.ls .Lskip_loop
|
||||
|
||||
.align 4
|
||||
.Loop_neon:
|
||||
////////////////////////////////////////////////////////////////
|
||||
// ((inp[0]*r^4+inp[2]*r^2+inp[4])*r^4+inp[6]*r^2
|
||||
// ((inp[1]*r^4+inp[3]*r^2+inp[5])*r^3+inp[7]*r
|
||||
// ___________________/
|
||||
// ((inp[0]*r^4+inp[2]*r^2+inp[4])*r^4+inp[6]*r^2+inp[8])*r^2
|
||||
// ((inp[1]*r^4+inp[3]*r^2+inp[5])*r^4+inp[7]*r^2+inp[9])*r
|
||||
// ___________________/ ____________________/
|
||||
//
|
||||
// Note that we start with inp[2:3]*r^2. This is because it
|
||||
// doesn't depend on reduction in previous iteration.
|
||||
////////////////////////////////////////////////////////////////
|
||||
// d4 = h0*r4 + h1*r3 + h2*r2 + h3*r1 + h4*r0
|
||||
// d3 = h0*r3 + h1*r2 + h2*r1 + h3*r0 + h4*5*r4
|
||||
// d2 = h0*r2 + h1*r1 + h2*r0 + h3*5*r4 + h4*5*r3
|
||||
// d1 = h0*r1 + h1*r0 + h2*5*r4 + h3*5*r3 + h4*5*r2
|
||||
// d0 = h0*r0 + h1*5*r4 + h2*5*r3 + h3*5*r2 + h4*5*r1
|
||||
|
||||
subs x2,x2,#64
|
||||
umull v23.2d,v14.2s,v7.s[2]
|
||||
csel x16,x17,x16,lo
|
||||
umull v22.2d,v14.2s,v5.s[2]
|
||||
umull v21.2d,v14.2s,v3.s[2]
|
||||
ldp x8,x12,[x16],#16 // inp[2:3] (or zero)
|
||||
umull v20.2d,v14.2s,v1.s[2]
|
||||
ldp x9,x13,[x16],#48
|
||||
umull v19.2d,v14.2s,v0.s[2]
|
||||
#ifdef __AARCH64EB__
|
||||
rev x8,x8
|
||||
rev x12,x12
|
||||
rev x9,x9
|
||||
rev x13,x13
|
||||
#endif
|
||||
|
||||
umlal v23.2d,v15.2s,v5.s[2]
|
||||
and x4,x8,#0x03ffffff // base 2^64 -> base 2^26
|
||||
umlal v22.2d,v15.2s,v3.s[2]
|
||||
and x5,x9,#0x03ffffff
|
||||
umlal v21.2d,v15.2s,v1.s[2]
|
||||
ubfx x6,x8,#26,#26
|
||||
umlal v20.2d,v15.2s,v0.s[2]
|
||||
ubfx x7,x9,#26,#26
|
||||
umlal v19.2d,v15.2s,v8.s[2]
|
||||
add x4,x4,x5,lsl#32 // bfi x4,x5,#32,#32
|
||||
|
||||
umlal v23.2d,v16.2s,v3.s[2]
|
||||
extr x8,x12,x8,#52
|
||||
umlal v22.2d,v16.2s,v1.s[2]
|
||||
extr x9,x13,x9,#52
|
||||
umlal v21.2d,v16.2s,v0.s[2]
|
||||
add x6,x6,x7,lsl#32 // bfi x6,x7,#32,#32
|
||||
umlal v20.2d,v16.2s,v8.s[2]
|
||||
fmov d14,x4
|
||||
umlal v19.2d,v16.2s,v6.s[2]
|
||||
and x8,x8,#0x03ffffff
|
||||
|
||||
umlal v23.2d,v17.2s,v1.s[2]
|
||||
and x9,x9,#0x03ffffff
|
||||
umlal v22.2d,v17.2s,v0.s[2]
|
||||
ubfx x10,x12,#14,#26
|
||||
umlal v21.2d,v17.2s,v8.s[2]
|
||||
ubfx x11,x13,#14,#26
|
||||
umlal v20.2d,v17.2s,v6.s[2]
|
||||
add x8,x8,x9,lsl#32 // bfi x8,x9,#32,#32
|
||||
umlal v19.2d,v17.2s,v4.s[2]
|
||||
fmov d15,x6
|
||||
|
||||
add v11.2s,v11.2s,v26.2s
|
||||
add x12,x3,x12,lsr#40
|
||||
umlal v23.2d,v18.2s,v0.s[2]
|
||||
add x13,x3,x13,lsr#40
|
||||
umlal v22.2d,v18.2s,v8.s[2]
|
||||
add x10,x10,x11,lsl#32 // bfi x10,x11,#32,#32
|
||||
umlal v21.2d,v18.2s,v6.s[2]
|
||||
add x12,x12,x13,lsl#32 // bfi x12,x13,#32,#32
|
||||
umlal v20.2d,v18.2s,v4.s[2]
|
||||
fmov d16,x8
|
||||
umlal v19.2d,v18.2s,v2.s[2]
|
||||
fmov d17,x10
|
||||
|
||||
////////////////////////////////////////////////////////////////
|
||||
// (hash+inp[0:1])*r^4 and accumulate
|
||||
|
||||
add v9.2s,v9.2s,v24.2s
|
||||
fmov d18,x12
|
||||
umlal v22.2d,v11.2s,v1.s[0]
|
||||
ldp x8,x12,[x1],#16 // inp[0:1]
|
||||
umlal v19.2d,v11.2s,v6.s[0]
|
||||
ldp x9,x13,[x1],#48
|
||||
umlal v23.2d,v11.2s,v3.s[0]
|
||||
umlal v20.2d,v11.2s,v8.s[0]
|
||||
umlal v21.2d,v11.2s,v0.s[0]
|
||||
#ifdef __AARCH64EB__
|
||||
rev x8,x8
|
||||
rev x12,x12
|
||||
rev x9,x9
|
||||
rev x13,x13
|
||||
#endif
|
||||
|
||||
add v10.2s,v10.2s,v25.2s
|
||||
umlal v22.2d,v9.2s,v5.s[0]
|
||||
umlal v23.2d,v9.2s,v7.s[0]
|
||||
and x4,x8,#0x03ffffff // base 2^64 -> base 2^26
|
||||
umlal v21.2d,v9.2s,v3.s[0]
|
||||
and x5,x9,#0x03ffffff
|
||||
umlal v19.2d,v9.2s,v0.s[0]
|
||||
ubfx x6,x8,#26,#26
|
||||
umlal v20.2d,v9.2s,v1.s[0]
|
||||
ubfx x7,x9,#26,#26
|
||||
|
||||
add v12.2s,v12.2s,v27.2s
|
||||
add x4,x4,x5,lsl#32 // bfi x4,x5,#32,#32
|
||||
umlal v22.2d,v10.2s,v3.s[0]
|
||||
extr x8,x12,x8,#52
|
||||
umlal v23.2d,v10.2s,v5.s[0]
|
||||
extr x9,x13,x9,#52
|
||||
umlal v19.2d,v10.2s,v8.s[0]
|
||||
add x6,x6,x7,lsl#32 // bfi x6,x7,#32,#32
|
||||
umlal v21.2d,v10.2s,v1.s[0]
|
||||
fmov d9,x4
|
||||
umlal v20.2d,v10.2s,v0.s[0]
|
||||
and x8,x8,#0x03ffffff
|
||||
|
||||
add v13.2s,v13.2s,v28.2s
|
||||
and x9,x9,#0x03ffffff
|
||||
umlal v22.2d,v12.2s,v0.s[0]
|
||||
ubfx x10,x12,#14,#26
|
||||
umlal v19.2d,v12.2s,v4.s[0]
|
||||
ubfx x11,x13,#14,#26
|
||||
umlal v23.2d,v12.2s,v1.s[0]
|
||||
add x8,x8,x9,lsl#32 // bfi x8,x9,#32,#32
|
||||
umlal v20.2d,v12.2s,v6.s[0]
|
||||
fmov d10,x6
|
||||
umlal v21.2d,v12.2s,v8.s[0]
|
||||
add x12,x3,x12,lsr#40
|
||||
|
||||
umlal v22.2d,v13.2s,v8.s[0]
|
||||
add x13,x3,x13,lsr#40
|
||||
umlal v19.2d,v13.2s,v2.s[0]
|
||||
add x10,x10,x11,lsl#32 // bfi x10,x11,#32,#32
|
||||
umlal v23.2d,v13.2s,v0.s[0]
|
||||
add x12,x12,x13,lsl#32 // bfi x12,x13,#32,#32
|
||||
umlal v20.2d,v13.2s,v4.s[0]
|
||||
fmov d11,x8
|
||||
umlal v21.2d,v13.2s,v6.s[0]
|
||||
fmov d12,x10
|
||||
fmov d13,x12
|
||||
|
||||
/////////////////////////////////////////////////////////////////
|
||||
// lazy reduction as discussed in "NEON crypto" by D.J. Bernstein
|
||||
// and P. Schwabe
|
||||
//
|
||||
// [see discussion in poly1305-armv4 module]
|
||||
|
||||
ushr v29.2d,v22.2d,#26
|
||||
xtn v27.2s,v22.2d
|
||||
ushr v30.2d,v19.2d,#26
|
||||
and v19.16b,v19.16b,v31.16b
|
||||
add v23.2d,v23.2d,v29.2d // h3 -> h4
|
||||
bic v27.2s,#0xfc,lsl#24 // &=0x03ffffff
|
||||
add v20.2d,v20.2d,v30.2d // h0 -> h1
|
||||
|
||||
ushr v29.2d,v23.2d,#26
|
||||
xtn v28.2s,v23.2d
|
||||
ushr v30.2d,v20.2d,#26
|
||||
xtn v25.2s,v20.2d
|
||||
bic v28.2s,#0xfc,lsl#24
|
||||
add v21.2d,v21.2d,v30.2d // h1 -> h2
|
||||
|
||||
add v19.2d,v19.2d,v29.2d
|
||||
shl v29.2d,v29.2d,#2
|
||||
shrn v30.2s,v21.2d,#26
|
||||
xtn v26.2s,v21.2d
|
||||
add v19.2d,v19.2d,v29.2d // h4 -> h0
|
||||
bic v25.2s,#0xfc,lsl#24
|
||||
add v27.2s,v27.2s,v30.2s // h2 -> h3
|
||||
bic v26.2s,#0xfc,lsl#24
|
||||
|
||||
shrn v29.2s,v19.2d,#26
|
||||
xtn v24.2s,v19.2d
|
||||
ushr v30.2s,v27.2s,#26
|
||||
bic v27.2s,#0xfc,lsl#24
|
||||
bic v24.2s,#0xfc,lsl#24
|
||||
add v25.2s,v25.2s,v29.2s // h0 -> h1
|
||||
add v28.2s,v28.2s,v30.2s // h3 -> h4
|
||||
|
||||
b.hi .Loop_neon
|
||||
|
||||
.Lskip_loop:
|
||||
dup v16.2d,v16.d[0]
|
||||
add v11.2s,v11.2s,v26.2s
|
||||
|
||||
////////////////////////////////////////////////////////////////
|
||||
// multiply (inp[0:1]+hash) or inp[2:3] by r^2:r^1
|
||||
|
||||
adds x2,x2,#32
|
||||
b.ne .Long_tail
|
||||
|
||||
dup v16.2d,v11.d[0]
|
||||
add v14.2s,v9.2s,v24.2s
|
||||
add v17.2s,v12.2s,v27.2s
|
||||
add v15.2s,v10.2s,v25.2s
|
||||
add v18.2s,v13.2s,v28.2s
|
||||
|
||||
.Long_tail:
|
||||
dup v14.2d,v14.d[0]
|
||||
umull2 v19.2d,v16.4s,v6.4s
|
||||
umull2 v22.2d,v16.4s,v1.4s
|
||||
umull2 v23.2d,v16.4s,v3.4s
|
||||
umull2 v21.2d,v16.4s,v0.4s
|
||||
umull2 v20.2d,v16.4s,v8.4s
|
||||
|
||||
dup v15.2d,v15.d[0]
|
||||
umlal2 v19.2d,v14.4s,v0.4s
|
||||
umlal2 v21.2d,v14.4s,v3.4s
|
||||
umlal2 v22.2d,v14.4s,v5.4s
|
||||
umlal2 v23.2d,v14.4s,v7.4s
|
||||
umlal2 v20.2d,v14.4s,v1.4s
|
||||
|
||||
dup v17.2d,v17.d[0]
|
||||
umlal2 v19.2d,v15.4s,v8.4s
|
||||
umlal2 v22.2d,v15.4s,v3.4s
|
||||
umlal2 v21.2d,v15.4s,v1.4s
|
||||
umlal2 v23.2d,v15.4s,v5.4s
|
||||
umlal2 v20.2d,v15.4s,v0.4s
|
||||
|
||||
dup v18.2d,v18.d[0]
|
||||
umlal2 v22.2d,v17.4s,v0.4s
|
||||
umlal2 v23.2d,v17.4s,v1.4s
|
||||
umlal2 v19.2d,v17.4s,v4.4s
|
||||
umlal2 v20.2d,v17.4s,v6.4s
|
||||
umlal2 v21.2d,v17.4s,v8.4s
|
||||
|
||||
umlal2 v22.2d,v18.4s,v8.4s
|
||||
umlal2 v19.2d,v18.4s,v2.4s
|
||||
umlal2 v23.2d,v18.4s,v0.4s
|
||||
umlal2 v20.2d,v18.4s,v4.4s
|
||||
umlal2 v21.2d,v18.4s,v6.4s
|
||||
|
||||
b.eq .Lshort_tail
|
||||
|
||||
////////////////////////////////////////////////////////////////
|
||||
// (hash+inp[0:1])*r^4:r^3 and accumulate
|
||||
|
||||
add v9.2s,v9.2s,v24.2s
|
||||
umlal v22.2d,v11.2s,v1.2s
|
||||
umlal v19.2d,v11.2s,v6.2s
|
||||
umlal v23.2d,v11.2s,v3.2s
|
||||
umlal v20.2d,v11.2s,v8.2s
|
||||
umlal v21.2d,v11.2s,v0.2s
|
||||
|
||||
add v10.2s,v10.2s,v25.2s
|
||||
umlal v22.2d,v9.2s,v5.2s
|
||||
umlal v19.2d,v9.2s,v0.2s
|
||||
umlal v23.2d,v9.2s,v7.2s
|
||||
umlal v20.2d,v9.2s,v1.2s
|
||||
umlal v21.2d,v9.2s,v3.2s
|
||||
|
||||
add v12.2s,v12.2s,v27.2s
|
||||
umlal v22.2d,v10.2s,v3.2s
|
||||
umlal v19.2d,v10.2s,v8.2s
|
||||
umlal v23.2d,v10.2s,v5.2s
|
||||
umlal v20.2d,v10.2s,v0.2s
|
||||
umlal v21.2d,v10.2s,v1.2s
|
||||
|
||||
add v13.2s,v13.2s,v28.2s
|
||||
umlal v22.2d,v12.2s,v0.2s
|
||||
umlal v19.2d,v12.2s,v4.2s
|
||||
umlal v23.2d,v12.2s,v1.2s
|
||||
umlal v20.2d,v12.2s,v6.2s
|
||||
umlal v21.2d,v12.2s,v8.2s
|
||||
|
||||
umlal v22.2d,v13.2s,v8.2s
|
||||
umlal v19.2d,v13.2s,v2.2s
|
||||
umlal v23.2d,v13.2s,v0.2s
|
||||
umlal v20.2d,v13.2s,v4.2s
|
||||
umlal v21.2d,v13.2s,v6.2s
|
||||
|
||||
.Lshort_tail:
|
||||
////////////////////////////////////////////////////////////////
|
||||
// horizontal add
|
||||
|
||||
addp v22.2d,v22.2d,v22.2d
|
||||
ldp d8,d9,[sp,#16] // meet ABI requirements
|
||||
addp v19.2d,v19.2d,v19.2d
|
||||
ldp d10,d11,[sp,#32]
|
||||
addp v23.2d,v23.2d,v23.2d
|
||||
ldp d12,d13,[sp,#48]
|
||||
addp v20.2d,v20.2d,v20.2d
|
||||
ldp d14,d15,[sp,#64]
|
||||
addp v21.2d,v21.2d,v21.2d
|
||||
ldr x30,[sp,#8]
|
||||
|
||||
////////////////////////////////////////////////////////////////
|
||||
// lazy reduction, but without narrowing
|
||||
|
||||
ushr v29.2d,v22.2d,#26
|
||||
and v22.16b,v22.16b,v31.16b
|
||||
ushr v30.2d,v19.2d,#26
|
||||
and v19.16b,v19.16b,v31.16b
|
||||
|
||||
add v23.2d,v23.2d,v29.2d // h3 -> h4
|
||||
add v20.2d,v20.2d,v30.2d // h0 -> h1
|
||||
|
||||
ushr v29.2d,v23.2d,#26
|
||||
and v23.16b,v23.16b,v31.16b
|
||||
ushr v30.2d,v20.2d,#26
|
||||
and v20.16b,v20.16b,v31.16b
|
||||
add v21.2d,v21.2d,v30.2d // h1 -> h2
|
||||
|
||||
add v19.2d,v19.2d,v29.2d
|
||||
shl v29.2d,v29.2d,#2
|
||||
ushr v30.2d,v21.2d,#26
|
||||
and v21.16b,v21.16b,v31.16b
|
||||
add v19.2d,v19.2d,v29.2d // h4 -> h0
|
||||
add v22.2d,v22.2d,v30.2d // h2 -> h3
|
||||
|
||||
ushr v29.2d,v19.2d,#26
|
||||
and v19.16b,v19.16b,v31.16b
|
||||
ushr v30.2d,v22.2d,#26
|
||||
and v22.16b,v22.16b,v31.16b
|
||||
add v20.2d,v20.2d,v29.2d // h0 -> h1
|
||||
add v23.2d,v23.2d,v30.2d // h3 -> h4
|
||||
|
||||
////////////////////////////////////////////////////////////////
|
||||
// write the result, can be partially reduced
|
||||
|
||||
st4 {v19.s,v20.s,v21.s,v22.s}[0],[x0],#16
|
||||
mov x4,#1
|
||||
st1 {v23.s}[0],[x0]
|
||||
str x4,[x0,#8] // set is_base2_26
|
||||
|
||||
ldr x29,[sp],#80
|
||||
.inst 0xd50323bf // autiasp
|
||||
ret
|
||||
.size poly1305_blocks_neon,.-poly1305_blocks_neon
|
||||
|
||||
.align 5
|
||||
.Lzeros:
|
||||
.long 0,0,0,0,0,0,0,0
|
||||
.asciz "Poly1305 for ARMv8, CRYPTOGAMS by @dot-asm"
|
||||
.align 2
|
||||
#if !defined(__KERNEL__) && !defined(_WIN64)
|
||||
.comm OPENSSL_armcap_P,4,4
|
||||
.hidden OPENSSL_armcap_P
|
||||
#endif
|
||||
231
arch/arm64/crypto/poly1305-glue.c
Normal file
231
arch/arm64/crypto/poly1305-glue.c
Normal file
|
|
@ -0,0 +1,231 @@
|
|||
// SPDX-License-Identifier: GPL-2.0
|
||||
/*
|
||||
* OpenSSL/Cryptogams accelerated Poly1305 transform for arm64
|
||||
*
|
||||
* Copyright (C) 2019 Linaro Ltd. <ard.biesheuvel@linaro.org>
|
||||
*/
|
||||
|
||||
#include <asm/hwcap.h>
|
||||
#include <asm/neon.h>
|
||||
#include <asm/simd.h>
|
||||
#include <asm/unaligned.h>
|
||||
#include <crypto/algapi.h>
|
||||
#include <crypto/internal/hash.h>
|
||||
#include <crypto/internal/poly1305.h>
|
||||
#include <crypto/internal/simd.h>
|
||||
#include <linux/cpufeature.h>
|
||||
#include <linux/crypto.h>
|
||||
#include <linux/jump_label.h>
|
||||
#include <linux/module.h>
|
||||
|
||||
asmlinkage void poly1305_init_arm64(void *state, const u8 *key);
|
||||
asmlinkage void poly1305_blocks(void *state, const u8 *src, u32 len, u32 hibit);
|
||||
asmlinkage void poly1305_blocks_neon(void *state, const u8 *src, u32 len, u32 hibit);
|
||||
asmlinkage void poly1305_emit(void *state, u8 *digest, const u32 *nonce);
|
||||
|
||||
static __ro_after_init DEFINE_STATIC_KEY_FALSE(have_neon);
|
||||
|
||||
void poly1305_init_arch(struct poly1305_desc_ctx *dctx, const u8 key[POLY1305_KEY_SIZE])
|
||||
{
|
||||
poly1305_init_arm64(&dctx->h, key);
|
||||
dctx->s[0] = get_unaligned_le32(key + 16);
|
||||
dctx->s[1] = get_unaligned_le32(key + 20);
|
||||
dctx->s[2] = get_unaligned_le32(key + 24);
|
||||
dctx->s[3] = get_unaligned_le32(key + 28);
|
||||
dctx->buflen = 0;
|
||||
}
|
||||
EXPORT_SYMBOL(poly1305_init_arch);
|
||||
|
||||
static int neon_poly1305_init(struct shash_desc *desc)
|
||||
{
|
||||
struct poly1305_desc_ctx *dctx = shash_desc_ctx(desc);
|
||||
|
||||
dctx->buflen = 0;
|
||||
dctx->rset = 0;
|
||||
dctx->sset = false;
|
||||
|
||||
return 0;
|
||||
}
|
||||
|
||||
static void neon_poly1305_blocks(struct poly1305_desc_ctx *dctx, const u8 *src,
|
||||
u32 len, u32 hibit, bool do_neon)
|
||||
{
|
||||
if (unlikely(!dctx->sset)) {
|
||||
if (!dctx->rset) {
|
||||
poly1305_init_arm64(&dctx->h, src);
|
||||
src += POLY1305_BLOCK_SIZE;
|
||||
len -= POLY1305_BLOCK_SIZE;
|
||||
dctx->rset = 1;
|
||||
}
|
||||
if (len >= POLY1305_BLOCK_SIZE) {
|
||||
dctx->s[0] = get_unaligned_le32(src + 0);
|
||||
dctx->s[1] = get_unaligned_le32(src + 4);
|
||||
dctx->s[2] = get_unaligned_le32(src + 8);
|
||||
dctx->s[3] = get_unaligned_le32(src + 12);
|
||||
src += POLY1305_BLOCK_SIZE;
|
||||
len -= POLY1305_BLOCK_SIZE;
|
||||
dctx->sset = true;
|
||||
}
|
||||
if (len < POLY1305_BLOCK_SIZE)
|
||||
return;
|
||||
}
|
||||
|
||||
len &= ~(POLY1305_BLOCK_SIZE - 1);
|
||||
|
||||
if (static_branch_likely(&have_neon) && likely(do_neon))
|
||||
poly1305_blocks_neon(&dctx->h, src, len, hibit);
|
||||
else
|
||||
poly1305_blocks(&dctx->h, src, len, hibit);
|
||||
}
|
||||
|
||||
static void neon_poly1305_do_update(struct poly1305_desc_ctx *dctx,
|
||||
const u8 *src, u32 len, bool do_neon)
|
||||
{
|
||||
if (unlikely(dctx->buflen)) {
|
||||
u32 bytes = min(len, POLY1305_BLOCK_SIZE - dctx->buflen);
|
||||
|
||||
memcpy(dctx->buf + dctx->buflen, src, bytes);
|
||||
src += bytes;
|
||||
len -= bytes;
|
||||
dctx->buflen += bytes;
|
||||
|
||||
if (dctx->buflen == POLY1305_BLOCK_SIZE) {
|
||||
neon_poly1305_blocks(dctx, dctx->buf,
|
||||
POLY1305_BLOCK_SIZE, 1, false);
|
||||
dctx->buflen = 0;
|
||||
}
|
||||
}
|
||||
|
||||
if (likely(len >= POLY1305_BLOCK_SIZE)) {
|
||||
neon_poly1305_blocks(dctx, src, len, 1, do_neon);
|
||||
src += round_down(len, POLY1305_BLOCK_SIZE);
|
||||
len %= POLY1305_BLOCK_SIZE;
|
||||
}
|
||||
|
||||
if (unlikely(len)) {
|
||||
dctx->buflen = len;
|
||||
memcpy(dctx->buf, src, len);
|
||||
}
|
||||
}
|
||||
|
||||
static int neon_poly1305_update(struct shash_desc *desc,
|
||||
const u8 *src, unsigned int srclen)
|
||||
{
|
||||
bool do_neon = crypto_simd_usable() && srclen > 128;
|
||||
struct poly1305_desc_ctx *dctx = shash_desc_ctx(desc);
|
||||
|
||||
if (static_branch_likely(&have_neon) && do_neon)
|
||||
kernel_neon_begin();
|
||||
neon_poly1305_do_update(dctx, src, srclen, do_neon);
|
||||
if (static_branch_likely(&have_neon) && do_neon)
|
||||
kernel_neon_end();
|
||||
return 0;
|
||||
}
|
||||
|
||||
void poly1305_update_arch(struct poly1305_desc_ctx *dctx, const u8 *src,
|
||||
unsigned int nbytes)
|
||||
{
|
||||
if (unlikely(dctx->buflen)) {
|
||||
u32 bytes = min(nbytes, POLY1305_BLOCK_SIZE - dctx->buflen);
|
||||
|
||||
memcpy(dctx->buf + dctx->buflen, src, bytes);
|
||||
src += bytes;
|
||||
nbytes -= bytes;
|
||||
dctx->buflen += bytes;
|
||||
|
||||
if (dctx->buflen == POLY1305_BLOCK_SIZE) {
|
||||
poly1305_blocks(&dctx->h, dctx->buf, POLY1305_BLOCK_SIZE, 1);
|
||||
dctx->buflen = 0;
|
||||
}
|
||||
}
|
||||
|
||||
if (likely(nbytes >= POLY1305_BLOCK_SIZE)) {
|
||||
unsigned int len = round_down(nbytes, POLY1305_BLOCK_SIZE);
|
||||
|
||||
if (static_branch_likely(&have_neon) && crypto_simd_usable()) {
|
||||
do {
|
||||
unsigned int todo = min_t(unsigned int, len, SZ_4K);
|
||||
|
||||
kernel_neon_begin();
|
||||
poly1305_blocks_neon(&dctx->h, src, todo, 1);
|
||||
kernel_neon_end();
|
||||
|
||||
len -= todo;
|
||||
src += todo;
|
||||
} while (len);
|
||||
} else {
|
||||
poly1305_blocks(&dctx->h, src, len, 1);
|
||||
src += len;
|
||||
}
|
||||
nbytes %= POLY1305_BLOCK_SIZE;
|
||||
}
|
||||
|
||||
if (unlikely(nbytes)) {
|
||||
dctx->buflen = nbytes;
|
||||
memcpy(dctx->buf, src, nbytes);
|
||||
}
|
||||
}
|
||||
EXPORT_SYMBOL(poly1305_update_arch);
|
||||
|
||||
void poly1305_final_arch(struct poly1305_desc_ctx *dctx, u8 *dst)
|
||||
{
|
||||
if (unlikely(dctx->buflen)) {
|
||||
dctx->buf[dctx->buflen++] = 1;
|
||||
memset(dctx->buf + dctx->buflen, 0,
|
||||
POLY1305_BLOCK_SIZE - dctx->buflen);
|
||||
poly1305_blocks(&dctx->h, dctx->buf, POLY1305_BLOCK_SIZE, 0);
|
||||
}
|
||||
|
||||
poly1305_emit(&dctx->h, dst, dctx->s);
|
||||
*dctx = (struct poly1305_desc_ctx){};
|
||||
}
|
||||
EXPORT_SYMBOL(poly1305_final_arch);
|
||||
|
||||
static int neon_poly1305_final(struct shash_desc *desc, u8 *dst)
|
||||
{
|
||||
struct poly1305_desc_ctx *dctx = shash_desc_ctx(desc);
|
||||
|
||||
if (unlikely(!dctx->sset))
|
||||
return -ENOKEY;
|
||||
|
||||
poly1305_final_arch(dctx, dst);
|
||||
return 0;
|
||||
}
|
||||
|
||||
static struct shash_alg neon_poly1305_alg = {
|
||||
.init = neon_poly1305_init,
|
||||
.update = neon_poly1305_update,
|
||||
.final = neon_poly1305_final,
|
||||
.digestsize = POLY1305_DIGEST_SIZE,
|
||||
.descsize = sizeof(struct poly1305_desc_ctx),
|
||||
|
||||
.base.cra_name = "poly1305",
|
||||
.base.cra_driver_name = "poly1305-neon",
|
||||
.base.cra_priority = 200,
|
||||
.base.cra_blocksize = POLY1305_BLOCK_SIZE,
|
||||
.base.cra_module = THIS_MODULE,
|
||||
};
|
||||
|
||||
static int __init neon_poly1305_mod_init(void)
|
||||
{
|
||||
if (!cpu_have_named_feature(ASIMD))
|
||||
return 0;
|
||||
|
||||
static_branch_enable(&have_neon);
|
||||
|
||||
return IS_REACHABLE(CONFIG_CRYPTO_HASH) ?
|
||||
crypto_register_shash(&neon_poly1305_alg) : 0;
|
||||
}
|
||||
|
||||
static void __exit neon_poly1305_mod_exit(void)
|
||||
{
|
||||
if (IS_REACHABLE(CONFIG_CRYPTO_HASH) && cpu_have_named_feature(ASIMD))
|
||||
crypto_unregister_shash(&neon_poly1305_alg);
|
||||
}
|
||||
|
||||
module_init(neon_poly1305_mod_init);
|
||||
module_exit(neon_poly1305_mod_exit);
|
||||
|
||||
MODULE_LICENSE("GPL v2");
|
||||
MODULE_ALIAS_CRYPTO("poly1305");
|
||||
MODULE_ALIAS_CRYPTO("poly1305-neon");
|
||||
|
|
@ -334,7 +334,7 @@ libs-$(CONFIG_MIPS_FP_SUPPORT) += arch/mips/math-emu/
|
|||
# See arch/mips/Kbuild for content of core part of the kernel
|
||||
core-y += arch/mips/
|
||||
|
||||
drivers-$(CONFIG_MIPS_CRC_SUPPORT) += arch/mips/crypto/
|
||||
drivers-y += arch/mips/crypto/
|
||||
drivers-$(CONFIG_OPROFILE) += arch/mips/oprofile/
|
||||
|
||||
# suspend and hibernation support
|
||||
|
|
|
|||
2
arch/mips/crypto/.gitignore
vendored
Normal file
2
arch/mips/crypto/.gitignore
vendored
Normal file
|
|
@ -0,0 +1,2 @@
|
|||
# SPDX-License-Identifier: GPL-2.0-only
|
||||
poly1305-core.S
|
||||
|
|
@ -4,3 +4,21 @@
|
|||
#
|
||||
|
||||
obj-$(CONFIG_CRYPTO_CRC32_MIPS) += crc32-mips.o
|
||||
|
||||
obj-$(CONFIG_CRYPTO_CHACHA_MIPS) += chacha-mips.o
|
||||
chacha-mips-y := chacha-core.o chacha-glue.o
|
||||
AFLAGS_chacha-core.o += -O2 # needed to fill branch delay slots
|
||||
|
||||
obj-$(CONFIG_CRYPTO_POLY1305_MIPS) += poly1305-mips.o
|
||||
poly1305-mips-y := poly1305-core.o poly1305-glue.o
|
||||
|
||||
perlasm-flavour-$(CONFIG_32BIT) := o32
|
||||
perlasm-flavour-$(CONFIG_64BIT) := 64
|
||||
|
||||
quiet_cmd_perlasm = PERLASM $@
|
||||
cmd_perlasm = $(PERL) $(<) $(perlasm-flavour-y) $(@)
|
||||
|
||||
$(obj)/poly1305-core.S: $(src)/poly1305-mips.pl FORCE
|
||||
$(call if_changed,perlasm)
|
||||
|
||||
targets += poly1305-core.S
|
||||
|
|
|
|||
497
arch/mips/crypto/chacha-core.S
Normal file
497
arch/mips/crypto/chacha-core.S
Normal file
|
|
@ -0,0 +1,497 @@
|
|||
/* SPDX-License-Identifier: GPL-2.0 OR MIT */
|
||||
/*
|
||||
* Copyright (C) 2016-2018 René van Dorst <opensource@vdorst.com>. All Rights Reserved.
|
||||
* Copyright (C) 2015-2019 Jason A. Donenfeld <Jason@zx2c4.com>. All Rights Reserved.
|
||||
*/
|
||||
|
||||
#define MASK_U32 0x3c
|
||||
#define CHACHA20_BLOCK_SIZE 64
|
||||
#define STACK_SIZE 32
|
||||
|
||||
#define X0 $t0
|
||||
#define X1 $t1
|
||||
#define X2 $t2
|
||||
#define X3 $t3
|
||||
#define X4 $t4
|
||||
#define X5 $t5
|
||||
#define X6 $t6
|
||||
#define X7 $t7
|
||||
#define X8 $t8
|
||||
#define X9 $t9
|
||||
#define X10 $v1
|
||||
#define X11 $s6
|
||||
#define X12 $s5
|
||||
#define X13 $s4
|
||||
#define X14 $s3
|
||||
#define X15 $s2
|
||||
/* Use regs which are overwritten on exit for Tx so we don't leak clear data. */
|
||||
#define T0 $s1
|
||||
#define T1 $s0
|
||||
#define T(n) T ## n
|
||||
#define X(n) X ## n
|
||||
|
||||
/* Input arguments */
|
||||
#define STATE $a0
|
||||
#define OUT $a1
|
||||
#define IN $a2
|
||||
#define BYTES $a3
|
||||
|
||||
/* Output argument */
|
||||
/* NONCE[0] is kept in a register and not in memory.
|
||||
* We don't want to touch original value in memory.
|
||||
* Must be incremented every loop iteration.
|
||||
*/
|
||||
#define NONCE_0 $v0
|
||||
|
||||
/* SAVED_X and SAVED_CA are set in the jump table.
|
||||
* Use regs which are overwritten on exit else we don't leak clear data.
|
||||
* They are used to handling the last bytes which are not multiple of 4.
|
||||
*/
|
||||
#define SAVED_X X15
|
||||
#define SAVED_CA $s7
|
||||
|
||||
#define IS_UNALIGNED $s7
|
||||
|
||||
#if __BYTE_ORDER__ == __ORDER_BIG_ENDIAN__
|
||||
#define MSB 0
|
||||
#define LSB 3
|
||||
#define ROTx rotl
|
||||
#define ROTR(n) rotr n, 24
|
||||
#define CPU_TO_LE32(n) \
|
||||
wsbh n; \
|
||||
rotr n, 16;
|
||||
#else
|
||||
#define MSB 3
|
||||
#define LSB 0
|
||||
#define ROTx rotr
|
||||
#define CPU_TO_LE32(n)
|
||||
#define ROTR(n)
|
||||
#endif
|
||||
|
||||
#define FOR_EACH_WORD(x) \
|
||||
x( 0); \
|
||||
x( 1); \
|
||||
x( 2); \
|
||||
x( 3); \
|
||||
x( 4); \
|
||||
x( 5); \
|
||||
x( 6); \
|
||||
x( 7); \
|
||||
x( 8); \
|
||||
x( 9); \
|
||||
x(10); \
|
||||
x(11); \
|
||||
x(12); \
|
||||
x(13); \
|
||||
x(14); \
|
||||
x(15);
|
||||
|
||||
#define FOR_EACH_WORD_REV(x) \
|
||||
x(15); \
|
||||
x(14); \
|
||||
x(13); \
|
||||
x(12); \
|
||||
x(11); \
|
||||
x(10); \
|
||||
x( 9); \
|
||||
x( 8); \
|
||||
x( 7); \
|
||||
x( 6); \
|
||||
x( 5); \
|
||||
x( 4); \
|
||||
x( 3); \
|
||||
x( 2); \
|
||||
x( 1); \
|
||||
x( 0);
|
||||
|
||||
#define PLUS_ONE_0 1
|
||||
#define PLUS_ONE_1 2
|
||||
#define PLUS_ONE_2 3
|
||||
#define PLUS_ONE_3 4
|
||||
#define PLUS_ONE_4 5
|
||||
#define PLUS_ONE_5 6
|
||||
#define PLUS_ONE_6 7
|
||||
#define PLUS_ONE_7 8
|
||||
#define PLUS_ONE_8 9
|
||||
#define PLUS_ONE_9 10
|
||||
#define PLUS_ONE_10 11
|
||||
#define PLUS_ONE_11 12
|
||||
#define PLUS_ONE_12 13
|
||||
#define PLUS_ONE_13 14
|
||||
#define PLUS_ONE_14 15
|
||||
#define PLUS_ONE_15 16
|
||||
#define PLUS_ONE(x) PLUS_ONE_ ## x
|
||||
#define _CONCAT3(a,b,c) a ## b ## c
|
||||
#define CONCAT3(a,b,c) _CONCAT3(a,b,c)
|
||||
|
||||
#define STORE_UNALIGNED(x) \
|
||||
CONCAT3(.Lchacha_mips_xor_unaligned_, PLUS_ONE(x), _b: ;) \
|
||||
.if (x != 12); \
|
||||
lw T0, (x*4)(STATE); \
|
||||
.endif; \
|
||||
lwl T1, (x*4)+MSB ## (IN); \
|
||||
lwr T1, (x*4)+LSB ## (IN); \
|
||||
.if (x == 12); \
|
||||
addu X ## x, NONCE_0; \
|
||||
.else; \
|
||||
addu X ## x, T0; \
|
||||
.endif; \
|
||||
CPU_TO_LE32(X ## x); \
|
||||
xor X ## x, T1; \
|
||||
swl X ## x, (x*4)+MSB ## (OUT); \
|
||||
swr X ## x, (x*4)+LSB ## (OUT);
|
||||
|
||||
#define STORE_ALIGNED(x) \
|
||||
CONCAT3(.Lchacha_mips_xor_aligned_, PLUS_ONE(x), _b: ;) \
|
||||
.if (x != 12); \
|
||||
lw T0, (x*4)(STATE); \
|
||||
.endif; \
|
||||
lw T1, (x*4) ## (IN); \
|
||||
.if (x == 12); \
|
||||
addu X ## x, NONCE_0; \
|
||||
.else; \
|
||||
addu X ## x, T0; \
|
||||
.endif; \
|
||||
CPU_TO_LE32(X ## x); \
|
||||
xor X ## x, T1; \
|
||||
sw X ## x, (x*4) ## (OUT);
|
||||
|
||||
/* Jump table macro.
|
||||
* Used for setup and handling the last bytes, which are not multiple of 4.
|
||||
* X15 is free to store Xn
|
||||
* Every jumptable entry must be equal in size.
|
||||
*/
|
||||
#define JMPTBL_ALIGNED(x) \
|
||||
.Lchacha_mips_jmptbl_aligned_ ## x: ; \
|
||||
.set noreorder; \
|
||||
b .Lchacha_mips_xor_aligned_ ## x ## _b; \
|
||||
.if (x == 12); \
|
||||
addu SAVED_X, X ## x, NONCE_0; \
|
||||
.else; \
|
||||
addu SAVED_X, X ## x, SAVED_CA; \
|
||||
.endif; \
|
||||
.set reorder
|
||||
|
||||
#define JMPTBL_UNALIGNED(x) \
|
||||
.Lchacha_mips_jmptbl_unaligned_ ## x: ; \
|
||||
.set noreorder; \
|
||||
b .Lchacha_mips_xor_unaligned_ ## x ## _b; \
|
||||
.if (x == 12); \
|
||||
addu SAVED_X, X ## x, NONCE_0; \
|
||||
.else; \
|
||||
addu SAVED_X, X ## x, SAVED_CA; \
|
||||
.endif; \
|
||||
.set reorder
|
||||
|
||||
#define AXR(A, B, C, D, K, L, M, N, V, W, Y, Z, S) \
|
||||
addu X(A), X(K); \
|
||||
addu X(B), X(L); \
|
||||
addu X(C), X(M); \
|
||||
addu X(D), X(N); \
|
||||
xor X(V), X(A); \
|
||||
xor X(W), X(B); \
|
||||
xor X(Y), X(C); \
|
||||
xor X(Z), X(D); \
|
||||
rotl X(V), S; \
|
||||
rotl X(W), S; \
|
||||
rotl X(Y), S; \
|
||||
rotl X(Z), S;
|
||||
|
||||
.text
|
||||
.set reorder
|
||||
.set noat
|
||||
.globl chacha_crypt_arch
|
||||
.ent chacha_crypt_arch
|
||||
chacha_crypt_arch:
|
||||
.frame $sp, STACK_SIZE, $ra
|
||||
|
||||
/* Load number of rounds */
|
||||
lw $at, 16($sp)
|
||||
|
||||
addiu $sp, -STACK_SIZE
|
||||
|
||||
/* Return bytes = 0. */
|
||||
beqz BYTES, .Lchacha_mips_end
|
||||
|
||||
lw NONCE_0, 48(STATE)
|
||||
|
||||
/* Save s0-s7 */
|
||||
sw $s0, 0($sp)
|
||||
sw $s1, 4($sp)
|
||||
sw $s2, 8($sp)
|
||||
sw $s3, 12($sp)
|
||||
sw $s4, 16($sp)
|
||||
sw $s5, 20($sp)
|
||||
sw $s6, 24($sp)
|
||||
sw $s7, 28($sp)
|
||||
|
||||
/* Test IN or OUT is unaligned.
|
||||
* IS_UNALIGNED = ( IN | OUT ) & 0x00000003
|
||||
*/
|
||||
or IS_UNALIGNED, IN, OUT
|
||||
andi IS_UNALIGNED, 0x3
|
||||
|
||||
b .Lchacha_rounds_start
|
||||
|
||||
.align 4
|
||||
.Loop_chacha_rounds:
|
||||
addiu IN, CHACHA20_BLOCK_SIZE
|
||||
addiu OUT, CHACHA20_BLOCK_SIZE
|
||||
addiu NONCE_0, 1
|
||||
|
||||
.Lchacha_rounds_start:
|
||||
lw X0, 0(STATE)
|
||||
lw X1, 4(STATE)
|
||||
lw X2, 8(STATE)
|
||||
lw X3, 12(STATE)
|
||||
|
||||
lw X4, 16(STATE)
|
||||
lw X5, 20(STATE)
|
||||
lw X6, 24(STATE)
|
||||
lw X7, 28(STATE)
|
||||
lw X8, 32(STATE)
|
||||
lw X9, 36(STATE)
|
||||
lw X10, 40(STATE)
|
||||
lw X11, 44(STATE)
|
||||
|
||||
move X12, NONCE_0
|
||||
lw X13, 52(STATE)
|
||||
lw X14, 56(STATE)
|
||||
lw X15, 60(STATE)
|
||||
|
||||
.Loop_chacha_xor_rounds:
|
||||
addiu $at, -2
|
||||
AXR( 0, 1, 2, 3, 4, 5, 6, 7, 12,13,14,15, 16);
|
||||
AXR( 8, 9,10,11, 12,13,14,15, 4, 5, 6, 7, 12);
|
||||
AXR( 0, 1, 2, 3, 4, 5, 6, 7, 12,13,14,15, 8);
|
||||
AXR( 8, 9,10,11, 12,13,14,15, 4, 5, 6, 7, 7);
|
||||
AXR( 0, 1, 2, 3, 5, 6, 7, 4, 15,12,13,14, 16);
|
||||
AXR(10,11, 8, 9, 15,12,13,14, 5, 6, 7, 4, 12);
|
||||
AXR( 0, 1, 2, 3, 5, 6, 7, 4, 15,12,13,14, 8);
|
||||
AXR(10,11, 8, 9, 15,12,13,14, 5, 6, 7, 4, 7);
|
||||
bnez $at, .Loop_chacha_xor_rounds
|
||||
|
||||
addiu BYTES, -(CHACHA20_BLOCK_SIZE)
|
||||
|
||||
/* Is data src/dst unaligned? Jump */
|
||||
bnez IS_UNALIGNED, .Loop_chacha_unaligned
|
||||
|
||||
/* Set number rounds here to fill delayslot. */
|
||||
lw $at, (STACK_SIZE+16)($sp)
|
||||
|
||||
/* BYTES < 0, it has no full block. */
|
||||
bltz BYTES, .Lchacha_mips_no_full_block_aligned
|
||||
|
||||
FOR_EACH_WORD_REV(STORE_ALIGNED)
|
||||
|
||||
/* BYTES > 0? Loop again. */
|
||||
bgtz BYTES, .Loop_chacha_rounds
|
||||
|
||||
/* Place this here to fill delay slot */
|
||||
addiu NONCE_0, 1
|
||||
|
||||
/* BYTES < 0? Handle last bytes */
|
||||
bltz BYTES, .Lchacha_mips_xor_bytes
|
||||
|
||||
.Lchacha_mips_xor_done:
|
||||
/* Restore used registers */
|
||||
lw $s0, 0($sp)
|
||||
lw $s1, 4($sp)
|
||||
lw $s2, 8($sp)
|
||||
lw $s3, 12($sp)
|
||||
lw $s4, 16($sp)
|
||||
lw $s5, 20($sp)
|
||||
lw $s6, 24($sp)
|
||||
lw $s7, 28($sp)
|
||||
|
||||
/* Write NONCE_0 back to right location in state */
|
||||
sw NONCE_0, 48(STATE)
|
||||
|
||||
.Lchacha_mips_end:
|
||||
addiu $sp, STACK_SIZE
|
||||
jr $ra
|
||||
|
||||
.Lchacha_mips_no_full_block_aligned:
|
||||
/* Restore the offset on BYTES */
|
||||
addiu BYTES, CHACHA20_BLOCK_SIZE
|
||||
|
||||
/* Get number of full WORDS */
|
||||
andi $at, BYTES, MASK_U32
|
||||
|
||||
/* Load upper half of jump table addr */
|
||||
lui T0, %hi(.Lchacha_mips_jmptbl_aligned_0)
|
||||
|
||||
/* Calculate lower half jump table offset */
|
||||
ins T0, $at, 1, 6
|
||||
|
||||
/* Add offset to STATE */
|
||||
addu T1, STATE, $at
|
||||
|
||||
/* Add lower half jump table addr */
|
||||
addiu T0, %lo(.Lchacha_mips_jmptbl_aligned_0)
|
||||
|
||||
/* Read value from STATE */
|
||||
lw SAVED_CA, 0(T1)
|
||||
|
||||
/* Store remaining bytecounter as negative value */
|
||||
subu BYTES, $at, BYTES
|
||||
|
||||
jr T0
|
||||
|
||||
/* Jump table */
|
||||
FOR_EACH_WORD(JMPTBL_ALIGNED)
|
||||
|
||||
|
||||
.Loop_chacha_unaligned:
|
||||
/* Set number rounds here to fill delayslot. */
|
||||
lw $at, (STACK_SIZE+16)($sp)
|
||||
|
||||
/* BYTES > 0, it has no full block. */
|
||||
bltz BYTES, .Lchacha_mips_no_full_block_unaligned
|
||||
|
||||
FOR_EACH_WORD_REV(STORE_UNALIGNED)
|
||||
|
||||
/* BYTES > 0? Loop again. */
|
||||
bgtz BYTES, .Loop_chacha_rounds
|
||||
|
||||
/* Write NONCE_0 back to right location in state */
|
||||
sw NONCE_0, 48(STATE)
|
||||
|
||||
.set noreorder
|
||||
/* Fall through to byte handling */
|
||||
bgez BYTES, .Lchacha_mips_xor_done
|
||||
.Lchacha_mips_xor_unaligned_0_b:
|
||||
.Lchacha_mips_xor_aligned_0_b:
|
||||
/* Place this here to fill delay slot */
|
||||
addiu NONCE_0, 1
|
||||
.set reorder
|
||||
|
||||
.Lchacha_mips_xor_bytes:
|
||||
addu IN, $at
|
||||
addu OUT, $at
|
||||
/* First byte */
|
||||
lbu T1, 0(IN)
|
||||
addiu $at, BYTES, 1
|
||||
CPU_TO_LE32(SAVED_X)
|
||||
ROTR(SAVED_X)
|
||||
xor T1, SAVED_X
|
||||
sb T1, 0(OUT)
|
||||
beqz $at, .Lchacha_mips_xor_done
|
||||
/* Second byte */
|
||||
lbu T1, 1(IN)
|
||||
addiu $at, BYTES, 2
|
||||
ROTx SAVED_X, 8
|
||||
xor T1, SAVED_X
|
||||
sb T1, 1(OUT)
|
||||
beqz $at, .Lchacha_mips_xor_done
|
||||
/* Third byte */
|
||||
lbu T1, 2(IN)
|
||||
ROTx SAVED_X, 8
|
||||
xor T1, SAVED_X
|
||||
sb T1, 2(OUT)
|
||||
b .Lchacha_mips_xor_done
|
||||
|
||||
.Lchacha_mips_no_full_block_unaligned:
|
||||
/* Restore the offset on BYTES */
|
||||
addiu BYTES, CHACHA20_BLOCK_SIZE
|
||||
|
||||
/* Get number of full WORDS */
|
||||
andi $at, BYTES, MASK_U32
|
||||
|
||||
/* Load upper half of jump table addr */
|
||||
lui T0, %hi(.Lchacha_mips_jmptbl_unaligned_0)
|
||||
|
||||
/* Calculate lower half jump table offset */
|
||||
ins T0, $at, 1, 6
|
||||
|
||||
/* Add offset to STATE */
|
||||
addu T1, STATE, $at
|
||||
|
||||
/* Add lower half jump table addr */
|
||||
addiu T0, %lo(.Lchacha_mips_jmptbl_unaligned_0)
|
||||
|
||||
/* Read value from STATE */
|
||||
lw SAVED_CA, 0(T1)
|
||||
|
||||
/* Store remaining bytecounter as negative value */
|
||||
subu BYTES, $at, BYTES
|
||||
|
||||
jr T0
|
||||
|
||||
/* Jump table */
|
||||
FOR_EACH_WORD(JMPTBL_UNALIGNED)
|
||||
.end chacha_crypt_arch
|
||||
.set at
|
||||
|
||||
/* Input arguments
|
||||
* STATE $a0
|
||||
* OUT $a1
|
||||
* NROUND $a2
|
||||
*/
|
||||
|
||||
#undef X12
|
||||
#undef X13
|
||||
#undef X14
|
||||
#undef X15
|
||||
|
||||
#define X12 $a3
|
||||
#define X13 $at
|
||||
#define X14 $v0
|
||||
#define X15 STATE
|
||||
|
||||
.set noat
|
||||
.globl hchacha_block_arch
|
||||
.ent hchacha_block_arch
|
||||
hchacha_block_arch:
|
||||
.frame $sp, STACK_SIZE, $ra
|
||||
|
||||
addiu $sp, -STACK_SIZE
|
||||
|
||||
/* Save X11(s6) */
|
||||
sw X11, 0($sp)
|
||||
|
||||
lw X0, 0(STATE)
|
||||
lw X1, 4(STATE)
|
||||
lw X2, 8(STATE)
|
||||
lw X3, 12(STATE)
|
||||
lw X4, 16(STATE)
|
||||
lw X5, 20(STATE)
|
||||
lw X6, 24(STATE)
|
||||
lw X7, 28(STATE)
|
||||
lw X8, 32(STATE)
|
||||
lw X9, 36(STATE)
|
||||
lw X10, 40(STATE)
|
||||
lw X11, 44(STATE)
|
||||
lw X12, 48(STATE)
|
||||
lw X13, 52(STATE)
|
||||
lw X14, 56(STATE)
|
||||
lw X15, 60(STATE)
|
||||
|
||||
.Loop_hchacha_xor_rounds:
|
||||
addiu $a2, -2
|
||||
AXR( 0, 1, 2, 3, 4, 5, 6, 7, 12,13,14,15, 16);
|
||||
AXR( 8, 9,10,11, 12,13,14,15, 4, 5, 6, 7, 12);
|
||||
AXR( 0, 1, 2, 3, 4, 5, 6, 7, 12,13,14,15, 8);
|
||||
AXR( 8, 9,10,11, 12,13,14,15, 4, 5, 6, 7, 7);
|
||||
AXR( 0, 1, 2, 3, 5, 6, 7, 4, 15,12,13,14, 16);
|
||||
AXR(10,11, 8, 9, 15,12,13,14, 5, 6, 7, 4, 12);
|
||||
AXR( 0, 1, 2, 3, 5, 6, 7, 4, 15,12,13,14, 8);
|
||||
AXR(10,11, 8, 9, 15,12,13,14, 5, 6, 7, 4, 7);
|
||||
bnez $a2, .Loop_hchacha_xor_rounds
|
||||
|
||||
/* Restore used register */
|
||||
lw X11, 0($sp)
|
||||
|
||||
sw X0, 0(OUT)
|
||||
sw X1, 4(OUT)
|
||||
sw X2, 8(OUT)
|
||||
sw X3, 12(OUT)
|
||||
sw X12, 16(OUT)
|
||||
sw X13, 20(OUT)
|
||||
sw X14, 24(OUT)
|
||||
sw X15, 28(OUT)
|
||||
|
||||
addiu $sp, STACK_SIZE
|
||||
jr $ra
|
||||
.end hchacha_block_arch
|
||||
.set at
|
||||
152
arch/mips/crypto/chacha-glue.c
Normal file
152
arch/mips/crypto/chacha-glue.c
Normal file
|
|
@ -0,0 +1,152 @@
|
|||
// SPDX-License-Identifier: GPL-2.0
|
||||
/*
|
||||
* MIPS accelerated ChaCha and XChaCha stream ciphers,
|
||||
* including ChaCha20 (RFC7539)
|
||||
*
|
||||
* Copyright (C) 2019 Linaro, Ltd. <ard.biesheuvel@linaro.org>
|
||||
*/
|
||||
|
||||
#include <asm/byteorder.h>
|
||||
#include <crypto/algapi.h>
|
||||
#include <crypto/internal/chacha.h>
|
||||
#include <crypto/internal/skcipher.h>
|
||||
#include <linux/kernel.h>
|
||||
#include <linux/module.h>
|
||||
|
||||
asmlinkage void chacha_crypt_arch(u32 *state, u8 *dst, const u8 *src,
|
||||
unsigned int bytes, int nrounds);
|
||||
EXPORT_SYMBOL(chacha_crypt_arch);
|
||||
|
||||
asmlinkage void hchacha_block_arch(const u32 *state, u32 *stream, int nrounds);
|
||||
EXPORT_SYMBOL(hchacha_block_arch);
|
||||
|
||||
void chacha_init_arch(u32 *state, const u32 *key, const u8 *iv)
|
||||
{
|
||||
chacha_init_generic(state, key, iv);
|
||||
}
|
||||
EXPORT_SYMBOL(chacha_init_arch);
|
||||
|
||||
static int chacha_mips_stream_xor(struct skcipher_request *req,
|
||||
const struct chacha_ctx *ctx, const u8 *iv)
|
||||
{
|
||||
struct skcipher_walk walk;
|
||||
u32 state[16];
|
||||
int err;
|
||||
|
||||
err = skcipher_walk_virt(&walk, req, false);
|
||||
|
||||
chacha_init_generic(state, ctx->key, iv);
|
||||
|
||||
while (walk.nbytes > 0) {
|
||||
unsigned int nbytes = walk.nbytes;
|
||||
|
||||
if (nbytes < walk.total)
|
||||
nbytes = round_down(nbytes, walk.stride);
|
||||
|
||||
chacha_crypt(state, walk.dst.virt.addr, walk.src.virt.addr,
|
||||
nbytes, ctx->nrounds);
|
||||
err = skcipher_walk_done(&walk, walk.nbytes - nbytes);
|
||||
}
|
||||
|
||||
return err;
|
||||
}
|
||||
|
||||
static int chacha_mips(struct skcipher_request *req)
|
||||
{
|
||||
struct crypto_skcipher *tfm = crypto_skcipher_reqtfm(req);
|
||||
struct chacha_ctx *ctx = crypto_skcipher_ctx(tfm);
|
||||
|
||||
return chacha_mips_stream_xor(req, ctx, req->iv);
|
||||
}
|
||||
|
||||
static int xchacha_mips(struct skcipher_request *req)
|
||||
{
|
||||
struct crypto_skcipher *tfm = crypto_skcipher_reqtfm(req);
|
||||
struct chacha_ctx *ctx = crypto_skcipher_ctx(tfm);
|
||||
struct chacha_ctx subctx;
|
||||
u32 state[16];
|
||||
u8 real_iv[16];
|
||||
|
||||
chacha_init_generic(state, ctx->key, req->iv);
|
||||
|
||||
hchacha_block(state, subctx.key, ctx->nrounds);
|
||||
subctx.nrounds = ctx->nrounds;
|
||||
|
||||
memcpy(&real_iv[0], req->iv + 24, 8);
|
||||
memcpy(&real_iv[8], req->iv + 16, 8);
|
||||
return chacha_mips_stream_xor(req, &subctx, real_iv);
|
||||
}
|
||||
|
||||
static struct skcipher_alg algs[] = {
|
||||
{
|
||||
.base.cra_name = "chacha20",
|
||||
.base.cra_driver_name = "chacha20-mips",
|
||||
.base.cra_priority = 200,
|
||||
.base.cra_blocksize = 1,
|
||||
.base.cra_ctxsize = sizeof(struct chacha_ctx),
|
||||
.base.cra_module = THIS_MODULE,
|
||||
|
||||
.min_keysize = CHACHA_KEY_SIZE,
|
||||
.max_keysize = CHACHA_KEY_SIZE,
|
||||
.ivsize = CHACHA_IV_SIZE,
|
||||
.chunksize = CHACHA_BLOCK_SIZE,
|
||||
.setkey = chacha20_setkey,
|
||||
.encrypt = chacha_mips,
|
||||
.decrypt = chacha_mips,
|
||||
}, {
|
||||
.base.cra_name = "xchacha20",
|
||||
.base.cra_driver_name = "xchacha20-mips",
|
||||
.base.cra_priority = 200,
|
||||
.base.cra_blocksize = 1,
|
||||
.base.cra_ctxsize = sizeof(struct chacha_ctx),
|
||||
.base.cra_module = THIS_MODULE,
|
||||
|
||||
.min_keysize = CHACHA_KEY_SIZE,
|
||||
.max_keysize = CHACHA_KEY_SIZE,
|
||||
.ivsize = XCHACHA_IV_SIZE,
|
||||
.chunksize = CHACHA_BLOCK_SIZE,
|
||||
.setkey = chacha20_setkey,
|
||||
.encrypt = xchacha_mips,
|
||||
.decrypt = xchacha_mips,
|
||||
}, {
|
||||
.base.cra_name = "xchacha12",
|
||||
.base.cra_driver_name = "xchacha12-mips",
|
||||
.base.cra_priority = 200,
|
||||
.base.cra_blocksize = 1,
|
||||
.base.cra_ctxsize = sizeof(struct chacha_ctx),
|
||||
.base.cra_module = THIS_MODULE,
|
||||
|
||||
.min_keysize = CHACHA_KEY_SIZE,
|
||||
.max_keysize = CHACHA_KEY_SIZE,
|
||||
.ivsize = XCHACHA_IV_SIZE,
|
||||
.chunksize = CHACHA_BLOCK_SIZE,
|
||||
.setkey = chacha12_setkey,
|
||||
.encrypt = xchacha_mips,
|
||||
.decrypt = xchacha_mips,
|
||||
}
|
||||
};
|
||||
|
||||
static int __init chacha_simd_mod_init(void)
|
||||
{
|
||||
return IS_REACHABLE(CONFIG_CRYPTO_BLKCIPHER) ?
|
||||
crypto_register_skciphers(algs, ARRAY_SIZE(algs)) : 0;
|
||||
}
|
||||
|
||||
static void __exit chacha_simd_mod_fini(void)
|
||||
{
|
||||
if (IS_REACHABLE(CONFIG_CRYPTO_BLKCIPHER))
|
||||
crypto_unregister_skciphers(algs, ARRAY_SIZE(algs));
|
||||
}
|
||||
|
||||
module_init(chacha_simd_mod_init);
|
||||
module_exit(chacha_simd_mod_fini);
|
||||
|
||||
MODULE_DESCRIPTION("ChaCha and XChaCha stream ciphers (MIPS accelerated)");
|
||||
MODULE_AUTHOR("Ard Biesheuvel <ard.biesheuvel@linaro.org>");
|
||||
MODULE_LICENSE("GPL v2");
|
||||
MODULE_ALIAS_CRYPTO("chacha20");
|
||||
MODULE_ALIAS_CRYPTO("chacha20-mips");
|
||||
MODULE_ALIAS_CRYPTO("xchacha20");
|
||||
MODULE_ALIAS_CRYPTO("xchacha20-mips");
|
||||
MODULE_ALIAS_CRYPTO("xchacha12");
|
||||
MODULE_ALIAS_CRYPTO("xchacha12-mips");
|
||||
191
arch/mips/crypto/poly1305-glue.c
Normal file
191
arch/mips/crypto/poly1305-glue.c
Normal file
|
|
@ -0,0 +1,191 @@
|
|||
// SPDX-License-Identifier: GPL-2.0
|
||||
/*
|
||||
* OpenSSL/Cryptogams accelerated Poly1305 transform for MIPS
|
||||
*
|
||||
* Copyright (C) 2019 Linaro Ltd. <ard.biesheuvel@linaro.org>
|
||||
*/
|
||||
|
||||
#include <asm/unaligned.h>
|
||||
#include <crypto/algapi.h>
|
||||
#include <crypto/internal/hash.h>
|
||||
#include <crypto/internal/poly1305.h>
|
||||
#include <linux/cpufeature.h>
|
||||
#include <linux/crypto.h>
|
||||
#include <linux/module.h>
|
||||
|
||||
asmlinkage void poly1305_init_mips(void *state, const u8 *key);
|
||||
asmlinkage void poly1305_blocks_mips(void *state, const u8 *src, u32 len, u32 hibit);
|
||||
asmlinkage void poly1305_emit_mips(void *state, u8 *digest, const u32 *nonce);
|
||||
|
||||
void poly1305_init_arch(struct poly1305_desc_ctx *dctx, const u8 key[POLY1305_KEY_SIZE])
|
||||
{
|
||||
poly1305_init_mips(&dctx->h, key);
|
||||
dctx->s[0] = get_unaligned_le32(key + 16);
|
||||
dctx->s[1] = get_unaligned_le32(key + 20);
|
||||
dctx->s[2] = get_unaligned_le32(key + 24);
|
||||
dctx->s[3] = get_unaligned_le32(key + 28);
|
||||
dctx->buflen = 0;
|
||||
}
|
||||
EXPORT_SYMBOL(poly1305_init_arch);
|
||||
|
||||
static int mips_poly1305_init(struct shash_desc *desc)
|
||||
{
|
||||
struct poly1305_desc_ctx *dctx = shash_desc_ctx(desc);
|
||||
|
||||
dctx->buflen = 0;
|
||||
dctx->rset = 0;
|
||||
dctx->sset = false;
|
||||
|
||||
return 0;
|
||||
}
|
||||
|
||||
static void mips_poly1305_blocks(struct poly1305_desc_ctx *dctx, const u8 *src,
|
||||
u32 len, u32 hibit)
|
||||
{
|
||||
if (unlikely(!dctx->sset)) {
|
||||
if (!dctx->rset) {
|
||||
poly1305_init_mips(&dctx->h, src);
|
||||
src += POLY1305_BLOCK_SIZE;
|
||||
len -= POLY1305_BLOCK_SIZE;
|
||||
dctx->rset = 1;
|
||||
}
|
||||
if (len >= POLY1305_BLOCK_SIZE) {
|
||||
dctx->s[0] = get_unaligned_le32(src + 0);
|
||||
dctx->s[1] = get_unaligned_le32(src + 4);
|
||||
dctx->s[2] = get_unaligned_le32(src + 8);
|
||||
dctx->s[3] = get_unaligned_le32(src + 12);
|
||||
src += POLY1305_BLOCK_SIZE;
|
||||
len -= POLY1305_BLOCK_SIZE;
|
||||
dctx->sset = true;
|
||||
}
|
||||
if (len < POLY1305_BLOCK_SIZE)
|
||||
return;
|
||||
}
|
||||
|
||||
len &= ~(POLY1305_BLOCK_SIZE - 1);
|
||||
|
||||
poly1305_blocks_mips(&dctx->h, src, len, hibit);
|
||||
}
|
||||
|
||||
static int mips_poly1305_update(struct shash_desc *desc, const u8 *src,
|
||||
unsigned int len)
|
||||
{
|
||||
struct poly1305_desc_ctx *dctx = shash_desc_ctx(desc);
|
||||
|
||||
if (unlikely(dctx->buflen)) {
|
||||
u32 bytes = min(len, POLY1305_BLOCK_SIZE - dctx->buflen);
|
||||
|
||||
memcpy(dctx->buf + dctx->buflen, src, bytes);
|
||||
src += bytes;
|
||||
len -= bytes;
|
||||
dctx->buflen += bytes;
|
||||
|
||||
if (dctx->buflen == POLY1305_BLOCK_SIZE) {
|
||||
mips_poly1305_blocks(dctx, dctx->buf, POLY1305_BLOCK_SIZE, 1);
|
||||
dctx->buflen = 0;
|
||||
}
|
||||
}
|
||||
|
||||
if (likely(len >= POLY1305_BLOCK_SIZE)) {
|
||||
mips_poly1305_blocks(dctx, src, len, 1);
|
||||
src += round_down(len, POLY1305_BLOCK_SIZE);
|
||||
len %= POLY1305_BLOCK_SIZE;
|
||||
}
|
||||
|
||||
if (unlikely(len)) {
|
||||
dctx->buflen = len;
|
||||
memcpy(dctx->buf, src, len);
|
||||
}
|
||||
return 0;
|
||||
}
|
||||
|
||||
void poly1305_update_arch(struct poly1305_desc_ctx *dctx, const u8 *src,
|
||||
unsigned int nbytes)
|
||||
{
|
||||
if (unlikely(dctx->buflen)) {
|
||||
u32 bytes = min(nbytes, POLY1305_BLOCK_SIZE - dctx->buflen);
|
||||
|
||||
memcpy(dctx->buf + dctx->buflen, src, bytes);
|
||||
src += bytes;
|
||||
nbytes -= bytes;
|
||||
dctx->buflen += bytes;
|
||||
|
||||
if (dctx->buflen == POLY1305_BLOCK_SIZE) {
|
||||
poly1305_blocks_mips(&dctx->h, dctx->buf,
|
||||
POLY1305_BLOCK_SIZE, 1);
|
||||
dctx->buflen = 0;
|
||||
}
|
||||
}
|
||||
|
||||
if (likely(nbytes >= POLY1305_BLOCK_SIZE)) {
|
||||
unsigned int len = round_down(nbytes, POLY1305_BLOCK_SIZE);
|
||||
|
||||
poly1305_blocks_mips(&dctx->h, src, len, 1);
|
||||
src += len;
|
||||
nbytes %= POLY1305_BLOCK_SIZE;
|
||||
}
|
||||
|
||||
if (unlikely(nbytes)) {
|
||||
dctx->buflen = nbytes;
|
||||
memcpy(dctx->buf, src, nbytes);
|
||||
}
|
||||
}
|
||||
EXPORT_SYMBOL(poly1305_update_arch);
|
||||
|
||||
void poly1305_final_arch(struct poly1305_desc_ctx *dctx, u8 *dst)
|
||||
{
|
||||
if (unlikely(dctx->buflen)) {
|
||||
dctx->buf[dctx->buflen++] = 1;
|
||||
memset(dctx->buf + dctx->buflen, 0,
|
||||
POLY1305_BLOCK_SIZE - dctx->buflen);
|
||||
poly1305_blocks_mips(&dctx->h, dctx->buf, POLY1305_BLOCK_SIZE, 0);
|
||||
}
|
||||
|
||||
poly1305_emit_mips(&dctx->h, dst, dctx->s);
|
||||
*dctx = (struct poly1305_desc_ctx){};
|
||||
}
|
||||
EXPORT_SYMBOL(poly1305_final_arch);
|
||||
|
||||
static int mips_poly1305_final(struct shash_desc *desc, u8 *dst)
|
||||
{
|
||||
struct poly1305_desc_ctx *dctx = shash_desc_ctx(desc);
|
||||
|
||||
if (unlikely(!dctx->sset))
|
||||
return -ENOKEY;
|
||||
|
||||
poly1305_final_arch(dctx, dst);
|
||||
return 0;
|
||||
}
|
||||
|
||||
static struct shash_alg mips_poly1305_alg = {
|
||||
.init = mips_poly1305_init,
|
||||
.update = mips_poly1305_update,
|
||||
.final = mips_poly1305_final,
|
||||
.digestsize = POLY1305_DIGEST_SIZE,
|
||||
.descsize = sizeof(struct poly1305_desc_ctx),
|
||||
|
||||
.base.cra_name = "poly1305",
|
||||
.base.cra_driver_name = "poly1305-mips",
|
||||
.base.cra_priority = 200,
|
||||
.base.cra_blocksize = POLY1305_BLOCK_SIZE,
|
||||
.base.cra_module = THIS_MODULE,
|
||||
};
|
||||
|
||||
static int __init mips_poly1305_mod_init(void)
|
||||
{
|
||||
return IS_REACHABLE(CONFIG_CRYPTO_HASH) ?
|
||||
crypto_register_shash(&mips_poly1305_alg) : 0;
|
||||
}
|
||||
|
||||
static void __exit mips_poly1305_mod_exit(void)
|
||||
{
|
||||
if (IS_REACHABLE(CONFIG_CRYPTO_HASH))
|
||||
crypto_unregister_shash(&mips_poly1305_alg);
|
||||
}
|
||||
|
||||
module_init(mips_poly1305_mod_init);
|
||||
module_exit(mips_poly1305_mod_exit);
|
||||
|
||||
MODULE_LICENSE("GPL v2");
|
||||
MODULE_ALIAS_CRYPTO("poly1305");
|
||||
MODULE_ALIAS_CRYPTO("poly1305-mips");
|
||||
1273
arch/mips/crypto/poly1305-mips.pl
Normal file
1273
arch/mips/crypto/poly1305-mips.pl
Normal file
File diff suppressed because it is too large
Load diff
|
|
@ -198,9 +198,10 @@ avx2_instr :=$(call as-instr,vpbroadcastb %xmm0$(comma)%ymm1,-DCONFIG_AS_AVX2=1)
|
|||
avx512_instr :=$(call as-instr,vpmovm2b %k1$(comma)%zmm5,-DCONFIG_AS_AVX512=1)
|
||||
sha1_ni_instr :=$(call as-instr,sha1msg1 %xmm0$(comma)%xmm1,-DCONFIG_AS_SHA1_NI=1)
|
||||
sha256_ni_instr :=$(call as-instr,sha256msg1 %xmm0$(comma)%xmm1,-DCONFIG_AS_SHA256_NI=1)
|
||||
adx_instr := $(call as-instr,adox %r10$(comma)%r10,-DCONFIG_AS_ADX=1)
|
||||
|
||||
KBUILD_AFLAGS += $(cfi) $(cfi-sigframe) $(cfi-sections) $(asinstr) $(avx_instr) $(avx2_instr) $(avx512_instr) $(sha1_ni_instr) $(sha256_ni_instr)
|
||||
KBUILD_CFLAGS += $(cfi) $(cfi-sigframe) $(cfi-sections) $(asinstr) $(avx_instr) $(avx2_instr) $(avx512_instr) $(sha1_ni_instr) $(sha256_ni_instr)
|
||||
KBUILD_AFLAGS += $(cfi) $(cfi-sigframe) $(cfi-sections) $(asinstr) $(avx_instr) $(avx2_instr) $(avx512_instr) $(sha1_ni_instr) $(sha256_ni_instr) $(adx_instr)
|
||||
KBUILD_CFLAGS += $(cfi) $(cfi-sigframe) $(cfi-sections) $(asinstr) $(avx_instr) $(avx2_instr) $(avx512_instr) $(sha1_ni_instr) $(sha256_ni_instr) $(adx_instr)
|
||||
|
||||
KBUILD_LDFLAGS := -m elf_$(UTS_MACHINE)
|
||||
|
||||
|
|
|
|||
|
|
@ -249,6 +249,7 @@ CONFIG_DM_VERITY_FEC=y
|
|||
CONFIG_DM_BOW=y
|
||||
CONFIG_NETDEVICES=y
|
||||
CONFIG_DUMMY=y
|
||||
CONFIG_WIREGUARD=y
|
||||
CONFIG_TUN=y
|
||||
CONFIG_VETH=y
|
||||
# CONFIG_ETHERNET is not set
|
||||
|
|
|
|||
1
arch/x86/crypto/.gitignore
vendored
Normal file
1
arch/x86/crypto/.gitignore
vendored
Normal file
|
|
@ -0,0 +1 @@
|
|||
poly1305-x86_64-cryptogams.S
|
||||
|
|
@ -11,6 +11,7 @@ avx2_supported := $(call as-instr,vpgatherdd %ymm0$(comma)(%eax$(comma)%ymm1\
|
|||
avx512_supported :=$(call as-instr,vpmovm2b %k1$(comma)%zmm5,yes,no)
|
||||
sha1_ni_supported :=$(call as-instr,sha1msg1 %xmm0$(comma)%xmm1,yes,no)
|
||||
sha256_ni_supported :=$(call as-instr,sha256msg1 %xmm0$(comma)%xmm1,yes,no)
|
||||
adx_supported := $(call as-instr,adox %r10$(comma)%r10,yes,no)
|
||||
|
||||
obj-$(CONFIG_CRYPTO_GLUE_HELPER_X86) += glue_helper.o
|
||||
|
||||
|
|
@ -40,6 +41,11 @@ obj-$(CONFIG_CRYPTO_AEGIS128_AESNI_SSE2) += aegis128-aesni.o
|
|||
obj-$(CONFIG_CRYPTO_NHPOLY1305_SSE2) += nhpoly1305-sse2.o
|
||||
obj-$(CONFIG_CRYPTO_NHPOLY1305_AVX2) += nhpoly1305-avx2.o
|
||||
|
||||
# These modules require the assembler to support ADX.
|
||||
ifeq ($(adx_supported),yes)
|
||||
obj-$(CONFIG_CRYPTO_CURVE25519_X86) += curve25519-x86_64.o
|
||||
endif
|
||||
|
||||
# These modules require assembler to support AVX.
|
||||
ifeq ($(avx_supported),yes)
|
||||
obj-$(CONFIG_CRYPTO_CAMELLIA_AESNI_AVX_X86_64) += \
|
||||
|
|
@ -74,6 +80,10 @@ nhpoly1305-sse2-y := nh-sse2-x86_64.o nhpoly1305-sse2-glue.o
|
|||
blake2s-x86_64-y := blake2s-shash.o
|
||||
obj-$(if $(CONFIG_CRYPTO_BLAKE2S_X86),y) += libblake2s-x86_64.o
|
||||
libblake2s-x86_64-y := blake2s-core.o blake2s-glue.o
|
||||
poly1305-x86_64-y := poly1305-x86_64-cryptogams.o poly1305_glue.o
|
||||
ifneq ($(CONFIG_CRYPTO_POLY1305_X86_64),)
|
||||
targets += poly1305-x86_64-cryptogams.S
|
||||
endif
|
||||
|
||||
ifeq ($(avx_supported),yes)
|
||||
camellia-aesni-avx-x86_64-y := camellia-aesni-avx-asm_64.o \
|
||||
|
|
@ -102,10 +112,8 @@ aesni-intel-y := aesni-intel_asm.o aesni-intel_glue.o
|
|||
aesni-intel-$(CONFIG_64BIT) += aesni-intel_avx-x86_64.o aes_ctrby8_avx-x86_64.o
|
||||
ghash-clmulni-intel-y := ghash-clmulni-intel_asm.o ghash-clmulni-intel_glue.o
|
||||
sha1-ssse3-y := sha1_ssse3_asm.o sha1_ssse3_glue.o
|
||||
poly1305-x86_64-y := poly1305-sse2-x86_64.o poly1305_glue.o
|
||||
ifeq ($(avx2_supported),yes)
|
||||
sha1-ssse3-y += sha1_avx2_x86_64_asm.o
|
||||
poly1305-x86_64-y += poly1305-avx2-x86_64.o
|
||||
endif
|
||||
ifeq ($(sha1_ni_supported),yes)
|
||||
sha1-ssse3-y += sha1_ni_asm.o
|
||||
|
|
@ -119,3 +127,8 @@ sha256-ssse3-y += sha256_ni_asm.o
|
|||
endif
|
||||
sha512-ssse3-y := sha512-ssse3-asm.o sha512-avx-asm.o sha512-avx2-asm.o sha512_ssse3_glue.o
|
||||
crct10dif-pclmul-y := crct10dif-pcl-asm_64.o crct10dif-pclmul_glue.o
|
||||
|
||||
quiet_cmd_perlasm = PERLASM $@
|
||||
cmd_perlasm = $(PERL) $< > $@
|
||||
$(obj)/%.S: $(src)/%.pl FORCE
|
||||
$(call if_changed,perlasm)
|
||||
|
|
|
|||
|
|
@ -69,7 +69,15 @@ static int __init blake2s_mod_init(void)
|
|||
XFEATURE_MASK_AVX512, NULL))
|
||||
static_branch_enable(&blake2s_use_avx512);
|
||||
|
||||
return 0;
|
||||
return IS_REACHABLE(CONFIG_CRYPTO_HASH) ?
|
||||
crypto_register_shashes(blake2s_algs,
|
||||
ARRAY_SIZE(blake2s_algs)) : 0;
|
||||
}
|
||||
|
||||
static void __exit blake2s_mod_exit(void)
|
||||
{
|
||||
if (IS_REACHABLE(CONFIG_CRYPTO_HASH) && boot_cpu_has(X86_FEATURE_SSSE3))
|
||||
crypto_unregister_shashes(blake2s_algs, ARRAY_SIZE(blake2s_algs));
|
||||
}
|
||||
|
||||
module_init(blake2s_mod_init);
|
||||
|
|
|
|||
|
|
@ -120,10 +120,10 @@ ENTRY(chacha_block_xor_ssse3)
|
|||
FRAME_BEGIN
|
||||
|
||||
# x0..3 = s0..3
|
||||
movdqa 0x00(%rdi),%xmm0
|
||||
movdqa 0x10(%rdi),%xmm1
|
||||
movdqa 0x20(%rdi),%xmm2
|
||||
movdqa 0x30(%rdi),%xmm3
|
||||
movdqu 0x00(%rdi),%xmm0
|
||||
movdqu 0x10(%rdi),%xmm1
|
||||
movdqu 0x20(%rdi),%xmm2
|
||||
movdqu 0x30(%rdi),%xmm3
|
||||
movdqa %xmm0,%xmm8
|
||||
movdqa %xmm1,%xmm9
|
||||
movdqa %xmm2,%xmm10
|
||||
|
|
@ -205,10 +205,10 @@ ENTRY(hchacha_block_ssse3)
|
|||
# %edx: nrounds
|
||||
FRAME_BEGIN
|
||||
|
||||
movdqa 0x00(%rdi),%xmm0
|
||||
movdqa 0x10(%rdi),%xmm1
|
||||
movdqa 0x20(%rdi),%xmm2
|
||||
movdqa 0x30(%rdi),%xmm3
|
||||
movdqu 0x00(%rdi),%xmm0
|
||||
movdqu 0x10(%rdi),%xmm1
|
||||
movdqu 0x20(%rdi),%xmm2
|
||||
movdqu 0x30(%rdi),%xmm3
|
||||
|
||||
mov %edx,%r8d
|
||||
call chacha_permute
|
||||
|
|
|
|||
|
|
@ -7,38 +7,36 @@
|
|||
*/
|
||||
|
||||
#include <crypto/algapi.h>
|
||||
#include <crypto/chacha.h>
|
||||
#include <crypto/internal/chacha.h>
|
||||
#include <crypto/internal/simd.h>
|
||||
#include <crypto/internal/skcipher.h>
|
||||
#include <linux/kernel.h>
|
||||
#include <linux/module.h>
|
||||
#include <asm/simd.h>
|
||||
|
||||
#define CHACHA_STATE_ALIGN 16
|
||||
|
||||
asmlinkage void chacha_block_xor_ssse3(u32 *state, u8 *dst, const u8 *src,
|
||||
unsigned int len, int nrounds);
|
||||
asmlinkage void chacha_4block_xor_ssse3(u32 *state, u8 *dst, const u8 *src,
|
||||
unsigned int len, int nrounds);
|
||||
asmlinkage void hchacha_block_ssse3(const u32 *state, u32 *out, int nrounds);
|
||||
#ifdef CONFIG_AS_AVX2
|
||||
|
||||
asmlinkage void chacha_2block_xor_avx2(u32 *state, u8 *dst, const u8 *src,
|
||||
unsigned int len, int nrounds);
|
||||
asmlinkage void chacha_4block_xor_avx2(u32 *state, u8 *dst, const u8 *src,
|
||||
unsigned int len, int nrounds);
|
||||
asmlinkage void chacha_8block_xor_avx2(u32 *state, u8 *dst, const u8 *src,
|
||||
unsigned int len, int nrounds);
|
||||
static bool chacha_use_avx2;
|
||||
#ifdef CONFIG_AS_AVX512
|
||||
|
||||
asmlinkage void chacha_2block_xor_avx512vl(u32 *state, u8 *dst, const u8 *src,
|
||||
unsigned int len, int nrounds);
|
||||
asmlinkage void chacha_4block_xor_avx512vl(u32 *state, u8 *dst, const u8 *src,
|
||||
unsigned int len, int nrounds);
|
||||
asmlinkage void chacha_8block_xor_avx512vl(u32 *state, u8 *dst, const u8 *src,
|
||||
unsigned int len, int nrounds);
|
||||
static bool chacha_use_avx512vl;
|
||||
#endif
|
||||
#endif
|
||||
|
||||
static __ro_after_init DEFINE_STATIC_KEY_FALSE(chacha_use_simd);
|
||||
static __ro_after_init DEFINE_STATIC_KEY_FALSE(chacha_use_avx2);
|
||||
static __ro_after_init DEFINE_STATIC_KEY_FALSE(chacha_use_avx512vl);
|
||||
|
||||
static unsigned int chacha_advance(unsigned int len, unsigned int maxblocks)
|
||||
{
|
||||
|
|
@ -49,9 +47,8 @@ static unsigned int chacha_advance(unsigned int len, unsigned int maxblocks)
|
|||
static void chacha_dosimd(u32 *state, u8 *dst, const u8 *src,
|
||||
unsigned int bytes, int nrounds)
|
||||
{
|
||||
#ifdef CONFIG_AS_AVX2
|
||||
#ifdef CONFIG_AS_AVX512
|
||||
if (chacha_use_avx512vl) {
|
||||
if (IS_ENABLED(CONFIG_AS_AVX512) &&
|
||||
static_branch_likely(&chacha_use_avx512vl)) {
|
||||
while (bytes >= CHACHA_BLOCK_SIZE * 8) {
|
||||
chacha_8block_xor_avx512vl(state, dst, src, bytes,
|
||||
nrounds);
|
||||
|
|
@ -79,8 +76,9 @@ static void chacha_dosimd(u32 *state, u8 *dst, const u8 *src,
|
|||
return;
|
||||
}
|
||||
}
|
||||
#endif
|
||||
if (chacha_use_avx2) {
|
||||
|
||||
if (IS_ENABLED(CONFIG_AS_AVX2) &&
|
||||
static_branch_likely(&chacha_use_avx2)) {
|
||||
while (bytes >= CHACHA_BLOCK_SIZE * 8) {
|
||||
chacha_8block_xor_avx2(state, dst, src, bytes, nrounds);
|
||||
bytes -= CHACHA_BLOCK_SIZE * 8;
|
||||
|
|
@ -104,7 +102,7 @@ static void chacha_dosimd(u32 *state, u8 *dst, const u8 *src,
|
|||
return;
|
||||
}
|
||||
}
|
||||
#endif
|
||||
|
||||
while (bytes >= CHACHA_BLOCK_SIZE * 4) {
|
||||
chacha_4block_xor_ssse3(state, dst, src, bytes, nrounds);
|
||||
bytes -= CHACHA_BLOCK_SIZE * 4;
|
||||
|
|
@ -123,37 +121,75 @@ static void chacha_dosimd(u32 *state, u8 *dst, const u8 *src,
|
|||
}
|
||||
}
|
||||
|
||||
static int chacha_simd_stream_xor(struct skcipher_walk *walk,
|
||||
void hchacha_block_arch(const u32 *state, u32 *stream, int nrounds)
|
||||
{
|
||||
if (!static_branch_likely(&chacha_use_simd) || !crypto_simd_usable()) {
|
||||
hchacha_block_generic(state, stream, nrounds);
|
||||
} else {
|
||||
kernel_fpu_begin();
|
||||
hchacha_block_ssse3(state, stream, nrounds);
|
||||
kernel_fpu_end();
|
||||
}
|
||||
}
|
||||
EXPORT_SYMBOL(hchacha_block_arch);
|
||||
|
||||
void chacha_init_arch(u32 *state, const u32 *key, const u8 *iv)
|
||||
{
|
||||
chacha_init_generic(state, key, iv);
|
||||
}
|
||||
EXPORT_SYMBOL(chacha_init_arch);
|
||||
|
||||
void chacha_crypt_arch(u32 *state, u8 *dst, const u8 *src, unsigned int bytes,
|
||||
int nrounds)
|
||||
{
|
||||
if (!static_branch_likely(&chacha_use_simd) || !crypto_simd_usable() ||
|
||||
bytes <= CHACHA_BLOCK_SIZE)
|
||||
return chacha_crypt_generic(state, dst, src, bytes, nrounds);
|
||||
|
||||
do {
|
||||
unsigned int todo = min_t(unsigned int, bytes, SZ_4K);
|
||||
|
||||
kernel_fpu_begin();
|
||||
chacha_dosimd(state, dst, src, todo, nrounds);
|
||||
kernel_fpu_end();
|
||||
|
||||
bytes -= todo;
|
||||
src += todo;
|
||||
dst += todo;
|
||||
} while (bytes);
|
||||
}
|
||||
EXPORT_SYMBOL(chacha_crypt_arch);
|
||||
|
||||
static int chacha_simd_stream_xor(struct skcipher_request *req,
|
||||
const struct chacha_ctx *ctx, const u8 *iv)
|
||||
{
|
||||
u32 *state, state_buf[16 + 2] __aligned(8);
|
||||
int next_yield = 4096; /* bytes until next FPU yield */
|
||||
int err = 0;
|
||||
u32 state[CHACHA_STATE_WORDS] __aligned(8);
|
||||
struct skcipher_walk walk;
|
||||
int err;
|
||||
|
||||
BUILD_BUG_ON(CHACHA_STATE_ALIGN != 16);
|
||||
state = PTR_ALIGN(state_buf + 0, CHACHA_STATE_ALIGN);
|
||||
err = skcipher_walk_virt(&walk, req, false);
|
||||
|
||||
crypto_chacha_init(state, ctx, iv);
|
||||
chacha_init_generic(state, ctx->key, iv);
|
||||
|
||||
while (walk->nbytes > 0) {
|
||||
unsigned int nbytes = walk->nbytes;
|
||||
while (walk.nbytes > 0) {
|
||||
unsigned int nbytes = walk.nbytes;
|
||||
|
||||
if (nbytes < walk->total) {
|
||||
nbytes = round_down(nbytes, walk->stride);
|
||||
next_yield -= nbytes;
|
||||
}
|
||||
if (nbytes < walk.total)
|
||||
nbytes = round_down(nbytes, walk.stride);
|
||||
|
||||
chacha_dosimd(state, walk->dst.virt.addr, walk->src.virt.addr,
|
||||
nbytes, ctx->nrounds);
|
||||
|
||||
if (next_yield <= 0) {
|
||||
/* temporarily allow preemption */
|
||||
kernel_fpu_end();
|
||||
if (!static_branch_likely(&chacha_use_simd) ||
|
||||
!crypto_simd_usable()) {
|
||||
chacha_crypt_generic(state, walk.dst.virt.addr,
|
||||
walk.src.virt.addr, nbytes,
|
||||
ctx->nrounds);
|
||||
} else {
|
||||
kernel_fpu_begin();
|
||||
next_yield = 4096;
|
||||
chacha_dosimd(state, walk.dst.virt.addr,
|
||||
walk.src.virt.addr, nbytes,
|
||||
ctx->nrounds);
|
||||
kernel_fpu_end();
|
||||
}
|
||||
|
||||
err = skcipher_walk_done(walk, walk->nbytes - nbytes);
|
||||
err = skcipher_walk_done(&walk, walk.nbytes - nbytes);
|
||||
}
|
||||
|
||||
return err;
|
||||
|
|
@ -163,55 +199,32 @@ static int chacha_simd(struct skcipher_request *req)
|
|||
{
|
||||
struct crypto_skcipher *tfm = crypto_skcipher_reqtfm(req);
|
||||
struct chacha_ctx *ctx = crypto_skcipher_ctx(tfm);
|
||||
struct skcipher_walk walk;
|
||||
int err;
|
||||
|
||||
if (req->cryptlen <= CHACHA_BLOCK_SIZE || !crypto_simd_usable())
|
||||
return crypto_chacha_crypt(req);
|
||||
|
||||
err = skcipher_walk_virt(&walk, req, true);
|
||||
if (err)
|
||||
return err;
|
||||
|
||||
kernel_fpu_begin();
|
||||
err = chacha_simd_stream_xor(&walk, ctx, req->iv);
|
||||
kernel_fpu_end();
|
||||
return err;
|
||||
return chacha_simd_stream_xor(req, ctx, req->iv);
|
||||
}
|
||||
|
||||
static int xchacha_simd(struct skcipher_request *req)
|
||||
{
|
||||
struct crypto_skcipher *tfm = crypto_skcipher_reqtfm(req);
|
||||
struct chacha_ctx *ctx = crypto_skcipher_ctx(tfm);
|
||||
struct skcipher_walk walk;
|
||||
u32 state[CHACHA_STATE_WORDS] __aligned(8);
|
||||
struct chacha_ctx subctx;
|
||||
u32 *state, state_buf[16 + 2] __aligned(8);
|
||||
u8 real_iv[16];
|
||||
int err;
|
||||
|
||||
if (req->cryptlen <= CHACHA_BLOCK_SIZE || !crypto_simd_usable())
|
||||
return crypto_xchacha_crypt(req);
|
||||
chacha_init_generic(state, ctx->key, req->iv);
|
||||
|
||||
err = skcipher_walk_virt(&walk, req, true);
|
||||
if (err)
|
||||
return err;
|
||||
|
||||
BUILD_BUG_ON(CHACHA_STATE_ALIGN != 16);
|
||||
state = PTR_ALIGN(state_buf + 0, CHACHA_STATE_ALIGN);
|
||||
crypto_chacha_init(state, ctx, req->iv);
|
||||
|
||||
kernel_fpu_begin();
|
||||
|
||||
hchacha_block_ssse3(state, subctx.key, ctx->nrounds);
|
||||
if (req->cryptlen > CHACHA_BLOCK_SIZE && crypto_simd_usable()) {
|
||||
kernel_fpu_begin();
|
||||
hchacha_block_ssse3(state, subctx.key, ctx->nrounds);
|
||||
kernel_fpu_end();
|
||||
} else {
|
||||
hchacha_block_generic(state, subctx.key, ctx->nrounds);
|
||||
}
|
||||
subctx.nrounds = ctx->nrounds;
|
||||
|
||||
memcpy(&real_iv[0], req->iv + 24, 8);
|
||||
memcpy(&real_iv[8], req->iv + 16, 8);
|
||||
err = chacha_simd_stream_xor(&walk, &subctx, real_iv);
|
||||
|
||||
kernel_fpu_end();
|
||||
|
||||
return err;
|
||||
return chacha_simd_stream_xor(req, &subctx, real_iv);
|
||||
}
|
||||
|
||||
static struct skcipher_alg algs[] = {
|
||||
|
|
@ -227,7 +240,7 @@ static struct skcipher_alg algs[] = {
|
|||
.max_keysize = CHACHA_KEY_SIZE,
|
||||
.ivsize = CHACHA_IV_SIZE,
|
||||
.chunksize = CHACHA_BLOCK_SIZE,
|
||||
.setkey = crypto_chacha20_setkey,
|
||||
.setkey = chacha20_setkey,
|
||||
.encrypt = chacha_simd,
|
||||
.decrypt = chacha_simd,
|
||||
}, {
|
||||
|
|
@ -242,7 +255,7 @@ static struct skcipher_alg algs[] = {
|
|||
.max_keysize = CHACHA_KEY_SIZE,
|
||||
.ivsize = XCHACHA_IV_SIZE,
|
||||
.chunksize = CHACHA_BLOCK_SIZE,
|
||||
.setkey = crypto_chacha20_setkey,
|
||||
.setkey = chacha20_setkey,
|
||||
.encrypt = xchacha_simd,
|
||||
.decrypt = xchacha_simd,
|
||||
}, {
|
||||
|
|
@ -257,7 +270,7 @@ static struct skcipher_alg algs[] = {
|
|||
.max_keysize = CHACHA_KEY_SIZE,
|
||||
.ivsize = XCHACHA_IV_SIZE,
|
||||
.chunksize = CHACHA_BLOCK_SIZE,
|
||||
.setkey = crypto_chacha12_setkey,
|
||||
.setkey = chacha12_setkey,
|
||||
.encrypt = xchacha_simd,
|
||||
.decrypt = xchacha_simd,
|
||||
},
|
||||
|
|
@ -266,24 +279,29 @@ static struct skcipher_alg algs[] = {
|
|||
static int __init chacha_simd_mod_init(void)
|
||||
{
|
||||
if (!boot_cpu_has(X86_FEATURE_SSSE3))
|
||||
return -ENODEV;
|
||||
return 0;
|
||||
|
||||
#ifdef CONFIG_AS_AVX2
|
||||
chacha_use_avx2 = boot_cpu_has(X86_FEATURE_AVX) &&
|
||||
boot_cpu_has(X86_FEATURE_AVX2) &&
|
||||
cpu_has_xfeatures(XFEATURE_MASK_SSE | XFEATURE_MASK_YMM, NULL);
|
||||
#ifdef CONFIG_AS_AVX512
|
||||
chacha_use_avx512vl = chacha_use_avx2 &&
|
||||
boot_cpu_has(X86_FEATURE_AVX512VL) &&
|
||||
boot_cpu_has(X86_FEATURE_AVX512BW); /* kmovq */
|
||||
#endif
|
||||
#endif
|
||||
return crypto_register_skciphers(algs, ARRAY_SIZE(algs));
|
||||
static_branch_enable(&chacha_use_simd);
|
||||
|
||||
if (IS_ENABLED(CONFIG_AS_AVX2) &&
|
||||
boot_cpu_has(X86_FEATURE_AVX) &&
|
||||
boot_cpu_has(X86_FEATURE_AVX2) &&
|
||||
cpu_has_xfeatures(XFEATURE_MASK_SSE | XFEATURE_MASK_YMM, NULL)) {
|
||||
static_branch_enable(&chacha_use_avx2);
|
||||
|
||||
if (IS_ENABLED(CONFIG_AS_AVX512) &&
|
||||
boot_cpu_has(X86_FEATURE_AVX512VL) &&
|
||||
boot_cpu_has(X86_FEATURE_AVX512BW)) /* kmovq */
|
||||
static_branch_enable(&chacha_use_avx512vl);
|
||||
}
|
||||
return IS_REACHABLE(CONFIG_CRYPTO_BLKCIPHER) ?
|
||||
crypto_register_skciphers(algs, ARRAY_SIZE(algs)) : 0;
|
||||
}
|
||||
|
||||
static void __exit chacha_simd_mod_fini(void)
|
||||
{
|
||||
crypto_unregister_skciphers(algs, ARRAY_SIZE(algs));
|
||||
if (IS_REACHABLE(CONFIG_CRYPTO_BLKCIPHER) && boot_cpu_has(X86_FEATURE_SSSE3))
|
||||
crypto_unregister_skciphers(algs, ARRAY_SIZE(algs));
|
||||
}
|
||||
|
||||
module_init(chacha_simd_mod_init);
|
||||
|
|
|
|||
1512
arch/x86/crypto/curve25519-x86_64.c
Normal file
1512
arch/x86/crypto/curve25519-x86_64.c
Normal file
File diff suppressed because it is too large
Load diff
|
|
@ -1,390 +0,0 @@
|
|||
/* SPDX-License-Identifier: GPL-2.0-or-later */
|
||||
/*
|
||||
* Poly1305 authenticator algorithm, RFC7539, x64 AVX2 functions
|
||||
*
|
||||
* Copyright (C) 2015 Martin Willi
|
||||
*/
|
||||
|
||||
#include <linux/linkage.h>
|
||||
|
||||
.section .rodata.cst32.ANMASK, "aM", @progbits, 32
|
||||
.align 32
|
||||
ANMASK: .octa 0x0000000003ffffff0000000003ffffff
|
||||
.octa 0x0000000003ffffff0000000003ffffff
|
||||
|
||||
.section .rodata.cst32.ORMASK, "aM", @progbits, 32
|
||||
.align 32
|
||||
ORMASK: .octa 0x00000000010000000000000001000000
|
||||
.octa 0x00000000010000000000000001000000
|
||||
|
||||
.text
|
||||
|
||||
#define h0 0x00(%rdi)
|
||||
#define h1 0x04(%rdi)
|
||||
#define h2 0x08(%rdi)
|
||||
#define h3 0x0c(%rdi)
|
||||
#define h4 0x10(%rdi)
|
||||
#define r0 0x00(%rdx)
|
||||
#define r1 0x04(%rdx)
|
||||
#define r2 0x08(%rdx)
|
||||
#define r3 0x0c(%rdx)
|
||||
#define r4 0x10(%rdx)
|
||||
#define u0 0x00(%r8)
|
||||
#define u1 0x04(%r8)
|
||||
#define u2 0x08(%r8)
|
||||
#define u3 0x0c(%r8)
|
||||
#define u4 0x10(%r8)
|
||||
#define w0 0x14(%r8)
|
||||
#define w1 0x18(%r8)
|
||||
#define w2 0x1c(%r8)
|
||||
#define w3 0x20(%r8)
|
||||
#define w4 0x24(%r8)
|
||||
#define y0 0x28(%r8)
|
||||
#define y1 0x2c(%r8)
|
||||
#define y2 0x30(%r8)
|
||||
#define y3 0x34(%r8)
|
||||
#define y4 0x38(%r8)
|
||||
#define m %rsi
|
||||
#define hc0 %ymm0
|
||||
#define hc1 %ymm1
|
||||
#define hc2 %ymm2
|
||||
#define hc3 %ymm3
|
||||
#define hc4 %ymm4
|
||||
#define hc0x %xmm0
|
||||
#define hc1x %xmm1
|
||||
#define hc2x %xmm2
|
||||
#define hc3x %xmm3
|
||||
#define hc4x %xmm4
|
||||
#define t1 %ymm5
|
||||
#define t2 %ymm6
|
||||
#define t1x %xmm5
|
||||
#define t2x %xmm6
|
||||
#define ruwy0 %ymm7
|
||||
#define ruwy1 %ymm8
|
||||
#define ruwy2 %ymm9
|
||||
#define ruwy3 %ymm10
|
||||
#define ruwy4 %ymm11
|
||||
#define ruwy0x %xmm7
|
||||
#define ruwy1x %xmm8
|
||||
#define ruwy2x %xmm9
|
||||
#define ruwy3x %xmm10
|
||||
#define ruwy4x %xmm11
|
||||
#define svxz1 %ymm12
|
||||
#define svxz2 %ymm13
|
||||
#define svxz3 %ymm14
|
||||
#define svxz4 %ymm15
|
||||
#define d0 %r9
|
||||
#define d1 %r10
|
||||
#define d2 %r11
|
||||
#define d3 %r12
|
||||
#define d4 %r13
|
||||
|
||||
ENTRY(poly1305_4block_avx2)
|
||||
# %rdi: Accumulator h[5]
|
||||
# %rsi: 64 byte input block m
|
||||
# %rdx: Poly1305 key r[5]
|
||||
# %rcx: Quadblock count
|
||||
# %r8: Poly1305 derived key r^2 u[5], r^3 w[5], r^4 y[5],
|
||||
|
||||
# This four-block variant uses loop unrolled block processing. It
|
||||
# requires 4 Poly1305 keys: r, r^2, r^3 and r^4:
|
||||
# h = (h + m) * r => h = (h + m1) * r^4 + m2 * r^3 + m3 * r^2 + m4 * r
|
||||
|
||||
vzeroupper
|
||||
push %rbx
|
||||
push %r12
|
||||
push %r13
|
||||
|
||||
# combine r0,u0,w0,y0
|
||||
vmovd y0,ruwy0x
|
||||
vmovd w0,t1x
|
||||
vpunpcklqdq t1,ruwy0,ruwy0
|
||||
vmovd u0,t1x
|
||||
vmovd r0,t2x
|
||||
vpunpcklqdq t2,t1,t1
|
||||
vperm2i128 $0x20,t1,ruwy0,ruwy0
|
||||
|
||||
# combine r1,u1,w1,y1 and s1=r1*5,v1=u1*5,x1=w1*5,z1=y1*5
|
||||
vmovd y1,ruwy1x
|
||||
vmovd w1,t1x
|
||||
vpunpcklqdq t1,ruwy1,ruwy1
|
||||
vmovd u1,t1x
|
||||
vmovd r1,t2x
|
||||
vpunpcklqdq t2,t1,t1
|
||||
vperm2i128 $0x20,t1,ruwy1,ruwy1
|
||||
vpslld $2,ruwy1,svxz1
|
||||
vpaddd ruwy1,svxz1,svxz1
|
||||
|
||||
# combine r2,u2,w2,y2 and s2=r2*5,v2=u2*5,x2=w2*5,z2=y2*5
|
||||
vmovd y2,ruwy2x
|
||||
vmovd w2,t1x
|
||||
vpunpcklqdq t1,ruwy2,ruwy2
|
||||
vmovd u2,t1x
|
||||
vmovd r2,t2x
|
||||
vpunpcklqdq t2,t1,t1
|
||||
vperm2i128 $0x20,t1,ruwy2,ruwy2
|
||||
vpslld $2,ruwy2,svxz2
|
||||
vpaddd ruwy2,svxz2,svxz2
|
||||
|
||||
# combine r3,u3,w3,y3 and s3=r3*5,v3=u3*5,x3=w3*5,z3=y3*5
|
||||
vmovd y3,ruwy3x
|
||||
vmovd w3,t1x
|
||||
vpunpcklqdq t1,ruwy3,ruwy3
|
||||
vmovd u3,t1x
|
||||
vmovd r3,t2x
|
||||
vpunpcklqdq t2,t1,t1
|
||||
vperm2i128 $0x20,t1,ruwy3,ruwy3
|
||||
vpslld $2,ruwy3,svxz3
|
||||
vpaddd ruwy3,svxz3,svxz3
|
||||
|
||||
# combine r4,u4,w4,y4 and s4=r4*5,v4=u4*5,x4=w4*5,z4=y4*5
|
||||
vmovd y4,ruwy4x
|
||||
vmovd w4,t1x
|
||||
vpunpcklqdq t1,ruwy4,ruwy4
|
||||
vmovd u4,t1x
|
||||
vmovd r4,t2x
|
||||
vpunpcklqdq t2,t1,t1
|
||||
vperm2i128 $0x20,t1,ruwy4,ruwy4
|
||||
vpslld $2,ruwy4,svxz4
|
||||
vpaddd ruwy4,svxz4,svxz4
|
||||
|
||||
.Ldoblock4:
|
||||
# hc0 = [m[48-51] & 0x3ffffff, m[32-35] & 0x3ffffff,
|
||||
# m[16-19] & 0x3ffffff, m[ 0- 3] & 0x3ffffff + h0]
|
||||
vmovd 0x00(m),hc0x
|
||||
vmovd 0x10(m),t1x
|
||||
vpunpcklqdq t1,hc0,hc0
|
||||
vmovd 0x20(m),t1x
|
||||
vmovd 0x30(m),t2x
|
||||
vpunpcklqdq t2,t1,t1
|
||||
vperm2i128 $0x20,t1,hc0,hc0
|
||||
vpand ANMASK(%rip),hc0,hc0
|
||||
vmovd h0,t1x
|
||||
vpaddd t1,hc0,hc0
|
||||
# hc1 = [(m[51-54] >> 2) & 0x3ffffff, (m[35-38] >> 2) & 0x3ffffff,
|
||||
# (m[19-22] >> 2) & 0x3ffffff, (m[ 3- 6] >> 2) & 0x3ffffff + h1]
|
||||
vmovd 0x03(m),hc1x
|
||||
vmovd 0x13(m),t1x
|
||||
vpunpcklqdq t1,hc1,hc1
|
||||
vmovd 0x23(m),t1x
|
||||
vmovd 0x33(m),t2x
|
||||
vpunpcklqdq t2,t1,t1
|
||||
vperm2i128 $0x20,t1,hc1,hc1
|
||||
vpsrld $2,hc1,hc1
|
||||
vpand ANMASK(%rip),hc1,hc1
|
||||
vmovd h1,t1x
|
||||
vpaddd t1,hc1,hc1
|
||||
# hc2 = [(m[54-57] >> 4) & 0x3ffffff, (m[38-41] >> 4) & 0x3ffffff,
|
||||
# (m[22-25] >> 4) & 0x3ffffff, (m[ 6- 9] >> 4) & 0x3ffffff + h2]
|
||||
vmovd 0x06(m),hc2x
|
||||
vmovd 0x16(m),t1x
|
||||
vpunpcklqdq t1,hc2,hc2
|
||||
vmovd 0x26(m),t1x
|
||||
vmovd 0x36(m),t2x
|
||||
vpunpcklqdq t2,t1,t1
|
||||
vperm2i128 $0x20,t1,hc2,hc2
|
||||
vpsrld $4,hc2,hc2
|
||||
vpand ANMASK(%rip),hc2,hc2
|
||||
vmovd h2,t1x
|
||||
vpaddd t1,hc2,hc2
|
||||
# hc3 = [(m[57-60] >> 6) & 0x3ffffff, (m[41-44] >> 6) & 0x3ffffff,
|
||||
# (m[25-28] >> 6) & 0x3ffffff, (m[ 9-12] >> 6) & 0x3ffffff + h3]
|
||||
vmovd 0x09(m),hc3x
|
||||
vmovd 0x19(m),t1x
|
||||
vpunpcklqdq t1,hc3,hc3
|
||||
vmovd 0x29(m),t1x
|
||||
vmovd 0x39(m),t2x
|
||||
vpunpcklqdq t2,t1,t1
|
||||
vperm2i128 $0x20,t1,hc3,hc3
|
||||
vpsrld $6,hc3,hc3
|
||||
vpand ANMASK(%rip),hc3,hc3
|
||||
vmovd h3,t1x
|
||||
vpaddd t1,hc3,hc3
|
||||
# hc4 = [(m[60-63] >> 8) | (1<<24), (m[44-47] >> 8) | (1<<24),
|
||||
# (m[28-31] >> 8) | (1<<24), (m[12-15] >> 8) | (1<<24) + h4]
|
||||
vmovd 0x0c(m),hc4x
|
||||
vmovd 0x1c(m),t1x
|
||||
vpunpcklqdq t1,hc4,hc4
|
||||
vmovd 0x2c(m),t1x
|
||||
vmovd 0x3c(m),t2x
|
||||
vpunpcklqdq t2,t1,t1
|
||||
vperm2i128 $0x20,t1,hc4,hc4
|
||||
vpsrld $8,hc4,hc4
|
||||
vpor ORMASK(%rip),hc4,hc4
|
||||
vmovd h4,t1x
|
||||
vpaddd t1,hc4,hc4
|
||||
|
||||
# t1 = [ hc0[3] * r0, hc0[2] * u0, hc0[1] * w0, hc0[0] * y0 ]
|
||||
vpmuludq hc0,ruwy0,t1
|
||||
# t1 += [ hc1[3] * s4, hc1[2] * v4, hc1[1] * x4, hc1[0] * z4 ]
|
||||
vpmuludq hc1,svxz4,t2
|
||||
vpaddq t2,t1,t1
|
||||
# t1 += [ hc2[3] * s3, hc2[2] * v3, hc2[1] * x3, hc2[0] * z3 ]
|
||||
vpmuludq hc2,svxz3,t2
|
||||
vpaddq t2,t1,t1
|
||||
# t1 += [ hc3[3] * s2, hc3[2] * v2, hc3[1] * x2, hc3[0] * z2 ]
|
||||
vpmuludq hc3,svxz2,t2
|
||||
vpaddq t2,t1,t1
|
||||
# t1 += [ hc4[3] * s1, hc4[2] * v1, hc4[1] * x1, hc4[0] * z1 ]
|
||||
vpmuludq hc4,svxz1,t2
|
||||
vpaddq t2,t1,t1
|
||||
# d0 = t1[0] + t1[1] + t[2] + t[3]
|
||||
vpermq $0xee,t1,t2
|
||||
vpaddq t2,t1,t1
|
||||
vpsrldq $8,t1,t2
|
||||
vpaddq t2,t1,t1
|
||||
vmovq t1x,d0
|
||||
|
||||
# t1 = [ hc0[3] * r1, hc0[2] * u1,hc0[1] * w1, hc0[0] * y1 ]
|
||||
vpmuludq hc0,ruwy1,t1
|
||||
# t1 += [ hc1[3] * r0, hc1[2] * u0, hc1[1] * w0, hc1[0] * y0 ]
|
||||
vpmuludq hc1,ruwy0,t2
|
||||
vpaddq t2,t1,t1
|
||||
# t1 += [ hc2[3] * s4, hc2[2] * v4, hc2[1] * x4, hc2[0] * z4 ]
|
||||
vpmuludq hc2,svxz4,t2
|
||||
vpaddq t2,t1,t1
|
||||
# t1 += [ hc3[3] * s3, hc3[2] * v3, hc3[1] * x3, hc3[0] * z3 ]
|
||||
vpmuludq hc3,svxz3,t2
|
||||
vpaddq t2,t1,t1
|
||||
# t1 += [ hc4[3] * s2, hc4[2] * v2, hc4[1] * x2, hc4[0] * z2 ]
|
||||
vpmuludq hc4,svxz2,t2
|
||||
vpaddq t2,t1,t1
|
||||
# d1 = t1[0] + t1[1] + t1[3] + t1[4]
|
||||
vpermq $0xee,t1,t2
|
||||
vpaddq t2,t1,t1
|
||||
vpsrldq $8,t1,t2
|
||||
vpaddq t2,t1,t1
|
||||
vmovq t1x,d1
|
||||
|
||||
# t1 = [ hc0[3] * r2, hc0[2] * u2, hc0[1] * w2, hc0[0] * y2 ]
|
||||
vpmuludq hc0,ruwy2,t1
|
||||
# t1 += [ hc1[3] * r1, hc1[2] * u1, hc1[1] * w1, hc1[0] * y1 ]
|
||||
vpmuludq hc1,ruwy1,t2
|
||||
vpaddq t2,t1,t1
|
||||
# t1 += [ hc2[3] * r0, hc2[2] * u0, hc2[1] * w0, hc2[0] * y0 ]
|
||||
vpmuludq hc2,ruwy0,t2
|
||||
vpaddq t2,t1,t1
|
||||
# t1 += [ hc3[3] * s4, hc3[2] * v4, hc3[1] * x4, hc3[0] * z4 ]
|
||||
vpmuludq hc3,svxz4,t2
|
||||
vpaddq t2,t1,t1
|
||||
# t1 += [ hc4[3] * s3, hc4[2] * v3, hc4[1] * x3, hc4[0] * z3 ]
|
||||
vpmuludq hc4,svxz3,t2
|
||||
vpaddq t2,t1,t1
|
||||
# d2 = t1[0] + t1[1] + t1[2] + t1[3]
|
||||
vpermq $0xee,t1,t2
|
||||
vpaddq t2,t1,t1
|
||||
vpsrldq $8,t1,t2
|
||||
vpaddq t2,t1,t1
|
||||
vmovq t1x,d2
|
||||
|
||||
# t1 = [ hc0[3] * r3, hc0[2] * u3, hc0[1] * w3, hc0[0] * y3 ]
|
||||
vpmuludq hc0,ruwy3,t1
|
||||
# t1 += [ hc1[3] * r2, hc1[2] * u2, hc1[1] * w2, hc1[0] * y2 ]
|
||||
vpmuludq hc1,ruwy2,t2
|
||||
vpaddq t2,t1,t1
|
||||
# t1 += [ hc2[3] * r1, hc2[2] * u1, hc2[1] * w1, hc2[0] * y1 ]
|
||||
vpmuludq hc2,ruwy1,t2
|
||||
vpaddq t2,t1,t1
|
||||
# t1 += [ hc3[3] * r0, hc3[2] * u0, hc3[1] * w0, hc3[0] * y0 ]
|
||||
vpmuludq hc3,ruwy0,t2
|
||||
vpaddq t2,t1,t1
|
||||
# t1 += [ hc4[3] * s4, hc4[2] * v4, hc4[1] * x4, hc4[0] * z4 ]
|
||||
vpmuludq hc4,svxz4,t2
|
||||
vpaddq t2,t1,t1
|
||||
# d3 = t1[0] + t1[1] + t1[2] + t1[3]
|
||||
vpermq $0xee,t1,t2
|
||||
vpaddq t2,t1,t1
|
||||
vpsrldq $8,t1,t2
|
||||
vpaddq t2,t1,t1
|
||||
vmovq t1x,d3
|
||||
|
||||
# t1 = [ hc0[3] * r4, hc0[2] * u4, hc0[1] * w4, hc0[0] * y4 ]
|
||||
vpmuludq hc0,ruwy4,t1
|
||||
# t1 += [ hc1[3] * r3, hc1[2] * u3, hc1[1] * w3, hc1[0] * y3 ]
|
||||
vpmuludq hc1,ruwy3,t2
|
||||
vpaddq t2,t1,t1
|
||||
# t1 += [ hc2[3] * r2, hc2[2] * u2, hc2[1] * w2, hc2[0] * y2 ]
|
||||
vpmuludq hc2,ruwy2,t2
|
||||
vpaddq t2,t1,t1
|
||||
# t1 += [ hc3[3] * r1, hc3[2] * u1, hc3[1] * w1, hc3[0] * y1 ]
|
||||
vpmuludq hc3,ruwy1,t2
|
||||
vpaddq t2,t1,t1
|
||||
# t1 += [ hc4[3] * r0, hc4[2] * u0, hc4[1] * w0, hc4[0] * y0 ]
|
||||
vpmuludq hc4,ruwy0,t2
|
||||
vpaddq t2,t1,t1
|
||||
# d4 = t1[0] + t1[1] + t1[2] + t1[3]
|
||||
vpermq $0xee,t1,t2
|
||||
vpaddq t2,t1,t1
|
||||
vpsrldq $8,t1,t2
|
||||
vpaddq t2,t1,t1
|
||||
vmovq t1x,d4
|
||||
|
||||
# Now do a partial reduction mod (2^130)-5, carrying h0 -> h1 -> h2 ->
|
||||
# h3 -> h4 -> h0 -> h1 to get h0,h2,h3,h4 < 2^26 and h1 < 2^26 + a small
|
||||
# amount. Careful: we must not assume the carry bits 'd0 >> 26',
|
||||
# 'd1 >> 26', 'd2 >> 26', 'd3 >> 26', and '(d4 >> 26) * 5' fit in 32-bit
|
||||
# integers. It's true in a single-block implementation, but not here.
|
||||
|
||||
# d1 += d0 >> 26
|
||||
mov d0,%rax
|
||||
shr $26,%rax
|
||||
add %rax,d1
|
||||
# h0 = d0 & 0x3ffffff
|
||||
mov d0,%rbx
|
||||
and $0x3ffffff,%ebx
|
||||
|
||||
# d2 += d1 >> 26
|
||||
mov d1,%rax
|
||||
shr $26,%rax
|
||||
add %rax,d2
|
||||
# h1 = d1 & 0x3ffffff
|
||||
mov d1,%rax
|
||||
and $0x3ffffff,%eax
|
||||
mov %eax,h1
|
||||
|
||||
# d3 += d2 >> 26
|
||||
mov d2,%rax
|
||||
shr $26,%rax
|
||||
add %rax,d3
|
||||
# h2 = d2 & 0x3ffffff
|
||||
mov d2,%rax
|
||||
and $0x3ffffff,%eax
|
||||
mov %eax,h2
|
||||
|
||||
# d4 += d3 >> 26
|
||||
mov d3,%rax
|
||||
shr $26,%rax
|
||||
add %rax,d4
|
||||
# h3 = d3 & 0x3ffffff
|
||||
mov d3,%rax
|
||||
and $0x3ffffff,%eax
|
||||
mov %eax,h3
|
||||
|
||||
# h0 += (d4 >> 26) * 5
|
||||
mov d4,%rax
|
||||
shr $26,%rax
|
||||
lea (%rax,%rax,4),%rax
|
||||
add %rax,%rbx
|
||||
# h4 = d4 & 0x3ffffff
|
||||
mov d4,%rax
|
||||
and $0x3ffffff,%eax
|
||||
mov %eax,h4
|
||||
|
||||
# h1 += h0 >> 26
|
||||
mov %rbx,%rax
|
||||
shr $26,%rax
|
||||
add %eax,h1
|
||||
# h0 = h0 & 0x3ffffff
|
||||
andl $0x3ffffff,%ebx
|
||||
mov %ebx,h0
|
||||
|
||||
add $0x40,m
|
||||
dec %rcx
|
||||
jnz .Ldoblock4
|
||||
|
||||
vzeroupper
|
||||
pop %r13
|
||||
pop %r12
|
||||
pop %rbx
|
||||
ret
|
||||
ENDPROC(poly1305_4block_avx2)
|
||||
|
|
@ -1,590 +0,0 @@
|
|||
/* SPDX-License-Identifier: GPL-2.0-or-later */
|
||||
/*
|
||||
* Poly1305 authenticator algorithm, RFC7539, x64 SSE2 functions
|
||||
*
|
||||
* Copyright (C) 2015 Martin Willi
|
||||
*/
|
||||
|
||||
#include <linux/linkage.h>
|
||||
|
||||
.section .rodata.cst16.ANMASK, "aM", @progbits, 16
|
||||
.align 16
|
||||
ANMASK: .octa 0x0000000003ffffff0000000003ffffff
|
||||
|
||||
.section .rodata.cst16.ORMASK, "aM", @progbits, 16
|
||||
.align 16
|
||||
ORMASK: .octa 0x00000000010000000000000001000000
|
||||
|
||||
.text
|
||||
|
||||
#define h0 0x00(%rdi)
|
||||
#define h1 0x04(%rdi)
|
||||
#define h2 0x08(%rdi)
|
||||
#define h3 0x0c(%rdi)
|
||||
#define h4 0x10(%rdi)
|
||||
#define r0 0x00(%rdx)
|
||||
#define r1 0x04(%rdx)
|
||||
#define r2 0x08(%rdx)
|
||||
#define r3 0x0c(%rdx)
|
||||
#define r4 0x10(%rdx)
|
||||
#define s1 0x00(%rsp)
|
||||
#define s2 0x04(%rsp)
|
||||
#define s3 0x08(%rsp)
|
||||
#define s4 0x0c(%rsp)
|
||||
#define m %rsi
|
||||
#define h01 %xmm0
|
||||
#define h23 %xmm1
|
||||
#define h44 %xmm2
|
||||
#define t1 %xmm3
|
||||
#define t2 %xmm4
|
||||
#define t3 %xmm5
|
||||
#define t4 %xmm6
|
||||
#define mask %xmm7
|
||||
#define d0 %r8
|
||||
#define d1 %r9
|
||||
#define d2 %r10
|
||||
#define d3 %r11
|
||||
#define d4 %r12
|
||||
|
||||
ENTRY(poly1305_block_sse2)
|
||||
# %rdi: Accumulator h[5]
|
||||
# %rsi: 16 byte input block m
|
||||
# %rdx: Poly1305 key r[5]
|
||||
# %rcx: Block count
|
||||
|
||||
# This single block variant tries to improve performance by doing two
|
||||
# multiplications in parallel using SSE instructions. There is quite
|
||||
# some quardword packing involved, hence the speedup is marginal.
|
||||
|
||||
push %rbx
|
||||
push %r12
|
||||
sub $0x10,%rsp
|
||||
|
||||
# s1..s4 = r1..r4 * 5
|
||||
mov r1,%eax
|
||||
lea (%eax,%eax,4),%eax
|
||||
mov %eax,s1
|
||||
mov r2,%eax
|
||||
lea (%eax,%eax,4),%eax
|
||||
mov %eax,s2
|
||||
mov r3,%eax
|
||||
lea (%eax,%eax,4),%eax
|
||||
mov %eax,s3
|
||||
mov r4,%eax
|
||||
lea (%eax,%eax,4),%eax
|
||||
mov %eax,s4
|
||||
|
||||
movdqa ANMASK(%rip),mask
|
||||
|
||||
.Ldoblock:
|
||||
# h01 = [0, h1, 0, h0]
|
||||
# h23 = [0, h3, 0, h2]
|
||||
# h44 = [0, h4, 0, h4]
|
||||
movd h0,h01
|
||||
movd h1,t1
|
||||
movd h2,h23
|
||||
movd h3,t2
|
||||
movd h4,h44
|
||||
punpcklqdq t1,h01
|
||||
punpcklqdq t2,h23
|
||||
punpcklqdq h44,h44
|
||||
|
||||
# h01 += [ (m[3-6] >> 2) & 0x3ffffff, m[0-3] & 0x3ffffff ]
|
||||
movd 0x00(m),t1
|
||||
movd 0x03(m),t2
|
||||
psrld $2,t2
|
||||
punpcklqdq t2,t1
|
||||
pand mask,t1
|
||||
paddd t1,h01
|
||||
# h23 += [ (m[9-12] >> 6) & 0x3ffffff, (m[6-9] >> 4) & 0x3ffffff ]
|
||||
movd 0x06(m),t1
|
||||
movd 0x09(m),t2
|
||||
psrld $4,t1
|
||||
psrld $6,t2
|
||||
punpcklqdq t2,t1
|
||||
pand mask,t1
|
||||
paddd t1,h23
|
||||
# h44 += [ (m[12-15] >> 8) | (1 << 24), (m[12-15] >> 8) | (1 << 24) ]
|
||||
mov 0x0c(m),%eax
|
||||
shr $8,%eax
|
||||
or $0x01000000,%eax
|
||||
movd %eax,t1
|
||||
pshufd $0xc4,t1,t1
|
||||
paddd t1,h44
|
||||
|
||||
# t1[0] = h0 * r0 + h2 * s3
|
||||
# t1[1] = h1 * s4 + h3 * s2
|
||||
movd r0,t1
|
||||
movd s4,t2
|
||||
punpcklqdq t2,t1
|
||||
pmuludq h01,t1
|
||||
movd s3,t2
|
||||
movd s2,t3
|
||||
punpcklqdq t3,t2
|
||||
pmuludq h23,t2
|
||||
paddq t2,t1
|
||||
# t2[0] = h0 * r1 + h2 * s4
|
||||
# t2[1] = h1 * r0 + h3 * s3
|
||||
movd r1,t2
|
||||
movd r0,t3
|
||||
punpcklqdq t3,t2
|
||||
pmuludq h01,t2
|
||||
movd s4,t3
|
||||
movd s3,t4
|
||||
punpcklqdq t4,t3
|
||||
pmuludq h23,t3
|
||||
paddq t3,t2
|
||||
# t3[0] = h4 * s1
|
||||
# t3[1] = h4 * s2
|
||||
movd s1,t3
|
||||
movd s2,t4
|
||||
punpcklqdq t4,t3
|
||||
pmuludq h44,t3
|
||||
# d0 = t1[0] + t1[1] + t3[0]
|
||||
# d1 = t2[0] + t2[1] + t3[1]
|
||||
movdqa t1,t4
|
||||
punpcklqdq t2,t4
|
||||
punpckhqdq t2,t1
|
||||
paddq t4,t1
|
||||
paddq t3,t1
|
||||
movq t1,d0
|
||||
psrldq $8,t1
|
||||
movq t1,d1
|
||||
|
||||
# t1[0] = h0 * r2 + h2 * r0
|
||||
# t1[1] = h1 * r1 + h3 * s4
|
||||
movd r2,t1
|
||||
movd r1,t2
|
||||
punpcklqdq t2,t1
|
||||
pmuludq h01,t1
|
||||
movd r0,t2
|
||||
movd s4,t3
|
||||
punpcklqdq t3,t2
|
||||
pmuludq h23,t2
|
||||
paddq t2,t1
|
||||
# t2[0] = h0 * r3 + h2 * r1
|
||||
# t2[1] = h1 * r2 + h3 * r0
|
||||
movd r3,t2
|
||||
movd r2,t3
|
||||
punpcklqdq t3,t2
|
||||
pmuludq h01,t2
|
||||
movd r1,t3
|
||||
movd r0,t4
|
||||
punpcklqdq t4,t3
|
||||
pmuludq h23,t3
|
||||
paddq t3,t2
|
||||
# t3[0] = h4 * s3
|
||||
# t3[1] = h4 * s4
|
||||
movd s3,t3
|
||||
movd s4,t4
|
||||
punpcklqdq t4,t3
|
||||
pmuludq h44,t3
|
||||
# d2 = t1[0] + t1[1] + t3[0]
|
||||
# d3 = t2[0] + t2[1] + t3[1]
|
||||
movdqa t1,t4
|
||||
punpcklqdq t2,t4
|
||||
punpckhqdq t2,t1
|
||||
paddq t4,t1
|
||||
paddq t3,t1
|
||||
movq t1,d2
|
||||
psrldq $8,t1
|
||||
movq t1,d3
|
||||
|
||||
# t1[0] = h0 * r4 + h2 * r2
|
||||
# t1[1] = h1 * r3 + h3 * r1
|
||||
movd r4,t1
|
||||
movd r3,t2
|
||||
punpcklqdq t2,t1
|
||||
pmuludq h01,t1
|
||||
movd r2,t2
|
||||
movd r1,t3
|
||||
punpcklqdq t3,t2
|
||||
pmuludq h23,t2
|
||||
paddq t2,t1
|
||||
# t3[0] = h4 * r0
|
||||
movd r0,t3
|
||||
pmuludq h44,t3
|
||||
# d4 = t1[0] + t1[1] + t3[0]
|
||||
movdqa t1,t4
|
||||
psrldq $8,t4
|
||||
paddq t4,t1
|
||||
paddq t3,t1
|
||||
movq t1,d4
|
||||
|
||||
# d1 += d0 >> 26
|
||||
mov d0,%rax
|
||||
shr $26,%rax
|
||||
add %rax,d1
|
||||
# h0 = d0 & 0x3ffffff
|
||||
mov d0,%rbx
|
||||
and $0x3ffffff,%ebx
|
||||
|
||||
# d2 += d1 >> 26
|
||||
mov d1,%rax
|
||||
shr $26,%rax
|
||||
add %rax,d2
|
||||
# h1 = d1 & 0x3ffffff
|
||||
mov d1,%rax
|
||||
and $0x3ffffff,%eax
|
||||
mov %eax,h1
|
||||
|
||||
# d3 += d2 >> 26
|
||||
mov d2,%rax
|
||||
shr $26,%rax
|
||||
add %rax,d3
|
||||
# h2 = d2 & 0x3ffffff
|
||||
mov d2,%rax
|
||||
and $0x3ffffff,%eax
|
||||
mov %eax,h2
|
||||
|
||||
# d4 += d3 >> 26
|
||||
mov d3,%rax
|
||||
shr $26,%rax
|
||||
add %rax,d4
|
||||
# h3 = d3 & 0x3ffffff
|
||||
mov d3,%rax
|
||||
and $0x3ffffff,%eax
|
||||
mov %eax,h3
|
||||
|
||||
# h0 += (d4 >> 26) * 5
|
||||
mov d4,%rax
|
||||
shr $26,%rax
|
||||
lea (%rax,%rax,4),%rax
|
||||
add %rax,%rbx
|
||||
# h4 = d4 & 0x3ffffff
|
||||
mov d4,%rax
|
||||
and $0x3ffffff,%eax
|
||||
mov %eax,h4
|
||||
|
||||
# h1 += h0 >> 26
|
||||
mov %rbx,%rax
|
||||
shr $26,%rax
|
||||
add %eax,h1
|
||||
# h0 = h0 & 0x3ffffff
|
||||
andl $0x3ffffff,%ebx
|
||||
mov %ebx,h0
|
||||
|
||||
add $0x10,m
|
||||
dec %rcx
|
||||
jnz .Ldoblock
|
||||
|
||||
# Zeroing of key material
|
||||
mov %rcx,0x00(%rsp)
|
||||
mov %rcx,0x08(%rsp)
|
||||
|
||||
add $0x10,%rsp
|
||||
pop %r12
|
||||
pop %rbx
|
||||
ret
|
||||
ENDPROC(poly1305_block_sse2)
|
||||
|
||||
|
||||
#define u0 0x00(%r8)
|
||||
#define u1 0x04(%r8)
|
||||
#define u2 0x08(%r8)
|
||||
#define u3 0x0c(%r8)
|
||||
#define u4 0x10(%r8)
|
||||
#define hc0 %xmm0
|
||||
#define hc1 %xmm1
|
||||
#define hc2 %xmm2
|
||||
#define hc3 %xmm5
|
||||
#define hc4 %xmm6
|
||||
#define ru0 %xmm7
|
||||
#define ru1 %xmm8
|
||||
#define ru2 %xmm9
|
||||
#define ru3 %xmm10
|
||||
#define ru4 %xmm11
|
||||
#define sv1 %xmm12
|
||||
#define sv2 %xmm13
|
||||
#define sv3 %xmm14
|
||||
#define sv4 %xmm15
|
||||
#undef d0
|
||||
#define d0 %r13
|
||||
|
||||
ENTRY(poly1305_2block_sse2)
|
||||
# %rdi: Accumulator h[5]
|
||||
# %rsi: 16 byte input block m
|
||||
# %rdx: Poly1305 key r[5]
|
||||
# %rcx: Doubleblock count
|
||||
# %r8: Poly1305 derived key r^2 u[5]
|
||||
|
||||
# This two-block variant further improves performance by using loop
|
||||
# unrolled block processing. This is more straight forward and does
|
||||
# less byte shuffling, but requires a second Poly1305 key r^2:
|
||||
# h = (h + m) * r => h = (h + m1) * r^2 + m2 * r
|
||||
|
||||
push %rbx
|
||||
push %r12
|
||||
push %r13
|
||||
|
||||
# combine r0,u0
|
||||
movd u0,ru0
|
||||
movd r0,t1
|
||||
punpcklqdq t1,ru0
|
||||
|
||||
# combine r1,u1 and s1=r1*5,v1=u1*5
|
||||
movd u1,ru1
|
||||
movd r1,t1
|
||||
punpcklqdq t1,ru1
|
||||
movdqa ru1,sv1
|
||||
pslld $2,sv1
|
||||
paddd ru1,sv1
|
||||
|
||||
# combine r2,u2 and s2=r2*5,v2=u2*5
|
||||
movd u2,ru2
|
||||
movd r2,t1
|
||||
punpcklqdq t1,ru2
|
||||
movdqa ru2,sv2
|
||||
pslld $2,sv2
|
||||
paddd ru2,sv2
|
||||
|
||||
# combine r3,u3 and s3=r3*5,v3=u3*5
|
||||
movd u3,ru3
|
||||
movd r3,t1
|
||||
punpcklqdq t1,ru3
|
||||
movdqa ru3,sv3
|
||||
pslld $2,sv3
|
||||
paddd ru3,sv3
|
||||
|
||||
# combine r4,u4 and s4=r4*5,v4=u4*5
|
||||
movd u4,ru4
|
||||
movd r4,t1
|
||||
punpcklqdq t1,ru4
|
||||
movdqa ru4,sv4
|
||||
pslld $2,sv4
|
||||
paddd ru4,sv4
|
||||
|
||||
.Ldoblock2:
|
||||
# hc0 = [ m[16-19] & 0x3ffffff, h0 + m[0-3] & 0x3ffffff ]
|
||||
movd 0x00(m),hc0
|
||||
movd 0x10(m),t1
|
||||
punpcklqdq t1,hc0
|
||||
pand ANMASK(%rip),hc0
|
||||
movd h0,t1
|
||||
paddd t1,hc0
|
||||
# hc1 = [ (m[19-22] >> 2) & 0x3ffffff, h1 + (m[3-6] >> 2) & 0x3ffffff ]
|
||||
movd 0x03(m),hc1
|
||||
movd 0x13(m),t1
|
||||
punpcklqdq t1,hc1
|
||||
psrld $2,hc1
|
||||
pand ANMASK(%rip),hc1
|
||||
movd h1,t1
|
||||
paddd t1,hc1
|
||||
# hc2 = [ (m[22-25] >> 4) & 0x3ffffff, h2 + (m[6-9] >> 4) & 0x3ffffff ]
|
||||
movd 0x06(m),hc2
|
||||
movd 0x16(m),t1
|
||||
punpcklqdq t1,hc2
|
||||
psrld $4,hc2
|
||||
pand ANMASK(%rip),hc2
|
||||
movd h2,t1
|
||||
paddd t1,hc2
|
||||
# hc3 = [ (m[25-28] >> 6) & 0x3ffffff, h3 + (m[9-12] >> 6) & 0x3ffffff ]
|
||||
movd 0x09(m),hc3
|
||||
movd 0x19(m),t1
|
||||
punpcklqdq t1,hc3
|
||||
psrld $6,hc3
|
||||
pand ANMASK(%rip),hc3
|
||||
movd h3,t1
|
||||
paddd t1,hc3
|
||||
# hc4 = [ (m[28-31] >> 8) | (1<<24), h4 + (m[12-15] >> 8) | (1<<24) ]
|
||||
movd 0x0c(m),hc4
|
||||
movd 0x1c(m),t1
|
||||
punpcklqdq t1,hc4
|
||||
psrld $8,hc4
|
||||
por ORMASK(%rip),hc4
|
||||
movd h4,t1
|
||||
paddd t1,hc4
|
||||
|
||||
# t1 = [ hc0[1] * r0, hc0[0] * u0 ]
|
||||
movdqa ru0,t1
|
||||
pmuludq hc0,t1
|
||||
# t1 += [ hc1[1] * s4, hc1[0] * v4 ]
|
||||
movdqa sv4,t2
|
||||
pmuludq hc1,t2
|
||||
paddq t2,t1
|
||||
# t1 += [ hc2[1] * s3, hc2[0] * v3 ]
|
||||
movdqa sv3,t2
|
||||
pmuludq hc2,t2
|
||||
paddq t2,t1
|
||||
# t1 += [ hc3[1] * s2, hc3[0] * v2 ]
|
||||
movdqa sv2,t2
|
||||
pmuludq hc3,t2
|
||||
paddq t2,t1
|
||||
# t1 += [ hc4[1] * s1, hc4[0] * v1 ]
|
||||
movdqa sv1,t2
|
||||
pmuludq hc4,t2
|
||||
paddq t2,t1
|
||||
# d0 = t1[0] + t1[1]
|
||||
movdqa t1,t2
|
||||
psrldq $8,t2
|
||||
paddq t2,t1
|
||||
movq t1,d0
|
||||
|
||||
# t1 = [ hc0[1] * r1, hc0[0] * u1 ]
|
||||
movdqa ru1,t1
|
||||
pmuludq hc0,t1
|
||||
# t1 += [ hc1[1] * r0, hc1[0] * u0 ]
|
||||
movdqa ru0,t2
|
||||
pmuludq hc1,t2
|
||||
paddq t2,t1
|
||||
# t1 += [ hc2[1] * s4, hc2[0] * v4 ]
|
||||
movdqa sv4,t2
|
||||
pmuludq hc2,t2
|
||||
paddq t2,t1
|
||||
# t1 += [ hc3[1] * s3, hc3[0] * v3 ]
|
||||
movdqa sv3,t2
|
||||
pmuludq hc3,t2
|
||||
paddq t2,t1
|
||||
# t1 += [ hc4[1] * s2, hc4[0] * v2 ]
|
||||
movdqa sv2,t2
|
||||
pmuludq hc4,t2
|
||||
paddq t2,t1
|
||||
# d1 = t1[0] + t1[1]
|
||||
movdqa t1,t2
|
||||
psrldq $8,t2
|
||||
paddq t2,t1
|
||||
movq t1,d1
|
||||
|
||||
# t1 = [ hc0[1] * r2, hc0[0] * u2 ]
|
||||
movdqa ru2,t1
|
||||
pmuludq hc0,t1
|
||||
# t1 += [ hc1[1] * r1, hc1[0] * u1 ]
|
||||
movdqa ru1,t2
|
||||
pmuludq hc1,t2
|
||||
paddq t2,t1
|
||||
# t1 += [ hc2[1] * r0, hc2[0] * u0 ]
|
||||
movdqa ru0,t2
|
||||
pmuludq hc2,t2
|
||||
paddq t2,t1
|
||||
# t1 += [ hc3[1] * s4, hc3[0] * v4 ]
|
||||
movdqa sv4,t2
|
||||
pmuludq hc3,t2
|
||||
paddq t2,t1
|
||||
# t1 += [ hc4[1] * s3, hc4[0] * v3 ]
|
||||
movdqa sv3,t2
|
||||
pmuludq hc4,t2
|
||||
paddq t2,t1
|
||||
# d2 = t1[0] + t1[1]
|
||||
movdqa t1,t2
|
||||
psrldq $8,t2
|
||||
paddq t2,t1
|
||||
movq t1,d2
|
||||
|
||||
# t1 = [ hc0[1] * r3, hc0[0] * u3 ]
|
||||
movdqa ru3,t1
|
||||
pmuludq hc0,t1
|
||||
# t1 += [ hc1[1] * r2, hc1[0] * u2 ]
|
||||
movdqa ru2,t2
|
||||
pmuludq hc1,t2
|
||||
paddq t2,t1
|
||||
# t1 += [ hc2[1] * r1, hc2[0] * u1 ]
|
||||
movdqa ru1,t2
|
||||
pmuludq hc2,t2
|
||||
paddq t2,t1
|
||||
# t1 += [ hc3[1] * r0, hc3[0] * u0 ]
|
||||
movdqa ru0,t2
|
||||
pmuludq hc3,t2
|
||||
paddq t2,t1
|
||||
# t1 += [ hc4[1] * s4, hc4[0] * v4 ]
|
||||
movdqa sv4,t2
|
||||
pmuludq hc4,t2
|
||||
paddq t2,t1
|
||||
# d3 = t1[0] + t1[1]
|
||||
movdqa t1,t2
|
||||
psrldq $8,t2
|
||||
paddq t2,t1
|
||||
movq t1,d3
|
||||
|
||||
# t1 = [ hc0[1] * r4, hc0[0] * u4 ]
|
||||
movdqa ru4,t1
|
||||
pmuludq hc0,t1
|
||||
# t1 += [ hc1[1] * r3, hc1[0] * u3 ]
|
||||
movdqa ru3,t2
|
||||
pmuludq hc1,t2
|
||||
paddq t2,t1
|
||||
# t1 += [ hc2[1] * r2, hc2[0] * u2 ]
|
||||
movdqa ru2,t2
|
||||
pmuludq hc2,t2
|
||||
paddq t2,t1
|
||||
# t1 += [ hc3[1] * r1, hc3[0] * u1 ]
|
||||
movdqa ru1,t2
|
||||
pmuludq hc3,t2
|
||||
paddq t2,t1
|
||||
# t1 += [ hc4[1] * r0, hc4[0] * u0 ]
|
||||
movdqa ru0,t2
|
||||
pmuludq hc4,t2
|
||||
paddq t2,t1
|
||||
# d4 = t1[0] + t1[1]
|
||||
movdqa t1,t2
|
||||
psrldq $8,t2
|
||||
paddq t2,t1
|
||||
movq t1,d4
|
||||
|
||||
# Now do a partial reduction mod (2^130)-5, carrying h0 -> h1 -> h2 ->
|
||||
# h3 -> h4 -> h0 -> h1 to get h0,h2,h3,h4 < 2^26 and h1 < 2^26 + a small
|
||||
# amount. Careful: we must not assume the carry bits 'd0 >> 26',
|
||||
# 'd1 >> 26', 'd2 >> 26', 'd3 >> 26', and '(d4 >> 26) * 5' fit in 32-bit
|
||||
# integers. It's true in a single-block implementation, but not here.
|
||||
|
||||
# d1 += d0 >> 26
|
||||
mov d0,%rax
|
||||
shr $26,%rax
|
||||
add %rax,d1
|
||||
# h0 = d0 & 0x3ffffff
|
||||
mov d0,%rbx
|
||||
and $0x3ffffff,%ebx
|
||||
|
||||
# d2 += d1 >> 26
|
||||
mov d1,%rax
|
||||
shr $26,%rax
|
||||
add %rax,d2
|
||||
# h1 = d1 & 0x3ffffff
|
||||
mov d1,%rax
|
||||
and $0x3ffffff,%eax
|
||||
mov %eax,h1
|
||||
|
||||
# d3 += d2 >> 26
|
||||
mov d2,%rax
|
||||
shr $26,%rax
|
||||
add %rax,d3
|
||||
# h2 = d2 & 0x3ffffff
|
||||
mov d2,%rax
|
||||
and $0x3ffffff,%eax
|
||||
mov %eax,h2
|
||||
|
||||
# d4 += d3 >> 26
|
||||
mov d3,%rax
|
||||
shr $26,%rax
|
||||
add %rax,d4
|
||||
# h3 = d3 & 0x3ffffff
|
||||
mov d3,%rax
|
||||
and $0x3ffffff,%eax
|
||||
mov %eax,h3
|
||||
|
||||
# h0 += (d4 >> 26) * 5
|
||||
mov d4,%rax
|
||||
shr $26,%rax
|
||||
lea (%rax,%rax,4),%rax
|
||||
add %rax,%rbx
|
||||
# h4 = d4 & 0x3ffffff
|
||||
mov d4,%rax
|
||||
and $0x3ffffff,%eax
|
||||
mov %eax,h4
|
||||
|
||||
# h1 += h0 >> 26
|
||||
mov %rbx,%rax
|
||||
shr $26,%rax
|
||||
add %eax,h1
|
||||
# h0 = h0 & 0x3ffffff
|
||||
andl $0x3ffffff,%ebx
|
||||
mov %ebx,h0
|
||||
|
||||
add $0x20,m
|
||||
dec %rcx
|
||||
jnz .Ldoblock2
|
||||
|
||||
pop %r13
|
||||
pop %r12
|
||||
pop %rbx
|
||||
ret
|
||||
ENDPROC(poly1305_2block_sse2)
|
||||
4265
arch/x86/crypto/poly1305-x86_64-cryptogams.pl
Normal file
4265
arch/x86/crypto/poly1305-x86_64-cryptogams.pl
Normal file
File diff suppressed because it is too large
Load diff
|
|
@ -1,131 +1,175 @@
|
|||
// SPDX-License-Identifier: GPL-2.0-or-later
|
||||
// SPDX-License-Identifier: GPL-2.0 OR MIT
|
||||
/*
|
||||
* Poly1305 authenticator algorithm, RFC7539, SIMD glue code
|
||||
*
|
||||
* Copyright (C) 2015 Martin Willi
|
||||
* Copyright (C) 2015-2019 Jason A. Donenfeld <Jason@zx2c4.com>. All Rights Reserved.
|
||||
*/
|
||||
|
||||
#include <crypto/algapi.h>
|
||||
#include <crypto/internal/hash.h>
|
||||
#include <crypto/internal/poly1305.h>
|
||||
#include <crypto/internal/simd.h>
|
||||
#include <crypto/poly1305.h>
|
||||
#include <linux/crypto.h>
|
||||
#include <linux/jump_label.h>
|
||||
#include <linux/kernel.h>
|
||||
#include <linux/module.h>
|
||||
#include <asm/intel-family.h>
|
||||
#include <asm/simd.h>
|
||||
|
||||
struct poly1305_simd_desc_ctx {
|
||||
struct poly1305_desc_ctx base;
|
||||
/* derived key u set? */
|
||||
bool uset;
|
||||
#ifdef CONFIG_AS_AVX2
|
||||
/* derived keys r^3, r^4 set? */
|
||||
bool wset;
|
||||
#endif
|
||||
/* derived Poly1305 key r^2 */
|
||||
u32 u[5];
|
||||
/* ... silently appended r^3 and r^4 when using AVX2 */
|
||||
asmlinkage void poly1305_init_x86_64(void *ctx,
|
||||
const u8 key[POLY1305_BLOCK_SIZE]);
|
||||
asmlinkage void poly1305_blocks_x86_64(void *ctx, const u8 *inp,
|
||||
const size_t len, const u32 padbit);
|
||||
asmlinkage void poly1305_emit_x86_64(void *ctx, u8 mac[POLY1305_DIGEST_SIZE],
|
||||
const u32 nonce[4]);
|
||||
asmlinkage void poly1305_emit_avx(void *ctx, u8 mac[POLY1305_DIGEST_SIZE],
|
||||
const u32 nonce[4]);
|
||||
asmlinkage void poly1305_blocks_avx(void *ctx, const u8 *inp, const size_t len,
|
||||
const u32 padbit);
|
||||
asmlinkage void poly1305_blocks_avx2(void *ctx, const u8 *inp, const size_t len,
|
||||
const u32 padbit);
|
||||
asmlinkage void poly1305_blocks_avx512(void *ctx, const u8 *inp,
|
||||
const size_t len, const u32 padbit);
|
||||
|
||||
static __ro_after_init DEFINE_STATIC_KEY_FALSE(poly1305_use_avx);
|
||||
static __ro_after_init DEFINE_STATIC_KEY_FALSE(poly1305_use_avx2);
|
||||
static __ro_after_init DEFINE_STATIC_KEY_FALSE(poly1305_use_avx512);
|
||||
|
||||
struct poly1305_arch_internal {
|
||||
union {
|
||||
struct {
|
||||
u32 h[5];
|
||||
u32 is_base2_26;
|
||||
};
|
||||
u64 hs[3];
|
||||
};
|
||||
u64 r[2];
|
||||
u64 pad;
|
||||
struct { u32 r2, r1, r4, r3; } rn[9];
|
||||
};
|
||||
|
||||
asmlinkage void poly1305_block_sse2(u32 *h, const u8 *src,
|
||||
const u32 *r, unsigned int blocks);
|
||||
asmlinkage void poly1305_2block_sse2(u32 *h, const u8 *src, const u32 *r,
|
||||
unsigned int blocks, const u32 *u);
|
||||
#ifdef CONFIG_AS_AVX2
|
||||
asmlinkage void poly1305_4block_avx2(u32 *h, const u8 *src, const u32 *r,
|
||||
unsigned int blocks, const u32 *u);
|
||||
static bool poly1305_use_avx2;
|
||||
#endif
|
||||
|
||||
static int poly1305_simd_init(struct shash_desc *desc)
|
||||
/* The AVX code uses base 2^26, while the scalar code uses base 2^64. If we hit
|
||||
* the unfortunate situation of using AVX and then having to go back to scalar
|
||||
* -- because the user is silly and has called the update function from two
|
||||
* separate contexts -- then we need to convert back to the original base before
|
||||
* proceeding. It is possible to reason that the initial reduction below is
|
||||
* sufficient given the implementation invariants. However, for an avoidance of
|
||||
* doubt and because this is not performance critical, we do the full reduction
|
||||
* anyway. Z3 proof of below function: https://xn--4db.cc/ltPtHCKN/py
|
||||
*/
|
||||
static void convert_to_base2_64(void *ctx)
|
||||
{
|
||||
struct poly1305_simd_desc_ctx *sctx = shash_desc_ctx(desc);
|
||||
struct poly1305_arch_internal *state = ctx;
|
||||
u32 cy;
|
||||
|
||||
sctx->uset = false;
|
||||
#ifdef CONFIG_AS_AVX2
|
||||
sctx->wset = false;
|
||||
#endif
|
||||
if (!state->is_base2_26)
|
||||
return;
|
||||
|
||||
return crypto_poly1305_init(desc);
|
||||
cy = state->h[0] >> 26; state->h[0] &= 0x3ffffff; state->h[1] += cy;
|
||||
cy = state->h[1] >> 26; state->h[1] &= 0x3ffffff; state->h[2] += cy;
|
||||
cy = state->h[2] >> 26; state->h[2] &= 0x3ffffff; state->h[3] += cy;
|
||||
cy = state->h[3] >> 26; state->h[3] &= 0x3ffffff; state->h[4] += cy;
|
||||
state->hs[0] = ((u64)state->h[2] << 52) | ((u64)state->h[1] << 26) | state->h[0];
|
||||
state->hs[1] = ((u64)state->h[4] << 40) | ((u64)state->h[3] << 14) | (state->h[2] >> 12);
|
||||
state->hs[2] = state->h[4] >> 24;
|
||||
#define ULT(a, b) ((a ^ ((a ^ b) | ((a - b) ^ b))) >> (sizeof(a) * 8 - 1))
|
||||
cy = (state->hs[2] >> 2) + (state->hs[2] & ~3ULL);
|
||||
state->hs[2] &= 3;
|
||||
state->hs[0] += cy;
|
||||
state->hs[1] += (cy = ULT(state->hs[0], cy));
|
||||
state->hs[2] += ULT(state->hs[1], cy);
|
||||
#undef ULT
|
||||
state->is_base2_26 = 0;
|
||||
}
|
||||
|
||||
static void poly1305_simd_mult(u32 *a, const u32 *b)
|
||||
static void poly1305_simd_init(void *ctx, const u8 key[POLY1305_BLOCK_SIZE])
|
||||
{
|
||||
u8 m[POLY1305_BLOCK_SIZE];
|
||||
|
||||
memset(m, 0, sizeof(m));
|
||||
/* The poly1305 block function adds a hi-bit to the accumulator which
|
||||
* we don't need for key multiplication; compensate for it. */
|
||||
a[4] -= 1 << 24;
|
||||
poly1305_block_sse2(a, m, b, 1);
|
||||
poly1305_init_x86_64(ctx, key);
|
||||
}
|
||||
|
||||
static unsigned int poly1305_simd_blocks(struct poly1305_desc_ctx *dctx,
|
||||
const u8 *src, unsigned int srclen)
|
||||
static void poly1305_simd_blocks(void *ctx, const u8 *inp, size_t len,
|
||||
const u32 padbit)
|
||||
{
|
||||
struct poly1305_simd_desc_ctx *sctx;
|
||||
unsigned int blocks, datalen;
|
||||
struct poly1305_arch_internal *state = ctx;
|
||||
|
||||
BUILD_BUG_ON(offsetof(struct poly1305_simd_desc_ctx, base));
|
||||
sctx = container_of(dctx, struct poly1305_simd_desc_ctx, base);
|
||||
/* SIMD disables preemption, so relax after processing each page. */
|
||||
BUILD_BUG_ON(SZ_4K < POLY1305_BLOCK_SIZE ||
|
||||
SZ_4K % POLY1305_BLOCK_SIZE);
|
||||
|
||||
if (!IS_ENABLED(CONFIG_AS_AVX) || !static_branch_likely(&poly1305_use_avx) ||
|
||||
(len < (POLY1305_BLOCK_SIZE * 18) && !state->is_base2_26) ||
|
||||
!crypto_simd_usable()) {
|
||||
convert_to_base2_64(ctx);
|
||||
poly1305_blocks_x86_64(ctx, inp, len, padbit);
|
||||
return;
|
||||
}
|
||||
|
||||
do {
|
||||
const size_t bytes = min_t(size_t, len, SZ_4K);
|
||||
|
||||
kernel_fpu_begin();
|
||||
if (IS_ENABLED(CONFIG_AS_AVX512) && static_branch_likely(&poly1305_use_avx512))
|
||||
poly1305_blocks_avx512(ctx, inp, bytes, padbit);
|
||||
else if (IS_ENABLED(CONFIG_AS_AVX2) && static_branch_likely(&poly1305_use_avx2))
|
||||
poly1305_blocks_avx2(ctx, inp, bytes, padbit);
|
||||
else
|
||||
poly1305_blocks_avx(ctx, inp, bytes, padbit);
|
||||
kernel_fpu_end();
|
||||
|
||||
len -= bytes;
|
||||
inp += bytes;
|
||||
} while (len);
|
||||
}
|
||||
|
||||
static void poly1305_simd_emit(void *ctx, u8 mac[POLY1305_DIGEST_SIZE],
|
||||
const u32 nonce[4])
|
||||
{
|
||||
if (!IS_ENABLED(CONFIG_AS_AVX) || !static_branch_likely(&poly1305_use_avx))
|
||||
poly1305_emit_x86_64(ctx, mac, nonce);
|
||||
else
|
||||
poly1305_emit_avx(ctx, mac, nonce);
|
||||
}
|
||||
|
||||
void poly1305_init_arch(struct poly1305_desc_ctx *dctx, const u8 key[POLY1305_KEY_SIZE])
|
||||
{
|
||||
poly1305_simd_init(&dctx->h, key);
|
||||
dctx->s[0] = get_unaligned_le32(&key[16]);
|
||||
dctx->s[1] = get_unaligned_le32(&key[20]);
|
||||
dctx->s[2] = get_unaligned_le32(&key[24]);
|
||||
dctx->s[3] = get_unaligned_le32(&key[28]);
|
||||
dctx->buflen = 0;
|
||||
dctx->sset = true;
|
||||
}
|
||||
EXPORT_SYMBOL(poly1305_init_arch);
|
||||
|
||||
static unsigned int crypto_poly1305_setdctxkey(struct poly1305_desc_ctx *dctx,
|
||||
const u8 *inp, unsigned int len)
|
||||
{
|
||||
unsigned int acc = 0;
|
||||
if (unlikely(!dctx->sset)) {
|
||||
datalen = crypto_poly1305_setdesckey(dctx, src, srclen);
|
||||
src += srclen - datalen;
|
||||
srclen = datalen;
|
||||
}
|
||||
|
||||
#ifdef CONFIG_AS_AVX2
|
||||
if (poly1305_use_avx2 && srclen >= POLY1305_BLOCK_SIZE * 4) {
|
||||
if (unlikely(!sctx->wset)) {
|
||||
if (!sctx->uset) {
|
||||
memcpy(sctx->u, dctx->r.r, sizeof(sctx->u));
|
||||
poly1305_simd_mult(sctx->u, dctx->r.r);
|
||||
sctx->uset = true;
|
||||
}
|
||||
memcpy(sctx->u + 5, sctx->u, sizeof(sctx->u));
|
||||
poly1305_simd_mult(sctx->u + 5, dctx->r.r);
|
||||
memcpy(sctx->u + 10, sctx->u + 5, sizeof(sctx->u));
|
||||
poly1305_simd_mult(sctx->u + 10, dctx->r.r);
|
||||
sctx->wset = true;
|
||||
if (!dctx->rset && len >= POLY1305_BLOCK_SIZE) {
|
||||
poly1305_simd_init(&dctx->h, inp);
|
||||
inp += POLY1305_BLOCK_SIZE;
|
||||
len -= POLY1305_BLOCK_SIZE;
|
||||
acc += POLY1305_BLOCK_SIZE;
|
||||
dctx->rset = 1;
|
||||
}
|
||||
blocks = srclen / (POLY1305_BLOCK_SIZE * 4);
|
||||
poly1305_4block_avx2(dctx->h.h, src, dctx->r.r, blocks,
|
||||
sctx->u);
|
||||
src += POLY1305_BLOCK_SIZE * 4 * blocks;
|
||||
srclen -= POLY1305_BLOCK_SIZE * 4 * blocks;
|
||||
}
|
||||
#endif
|
||||
if (likely(srclen >= POLY1305_BLOCK_SIZE * 2)) {
|
||||
if (unlikely(!sctx->uset)) {
|
||||
memcpy(sctx->u, dctx->r.r, sizeof(sctx->u));
|
||||
poly1305_simd_mult(sctx->u, dctx->r.r);
|
||||
sctx->uset = true;
|
||||
if (len >= POLY1305_BLOCK_SIZE) {
|
||||
dctx->s[0] = get_unaligned_le32(&inp[0]);
|
||||
dctx->s[1] = get_unaligned_le32(&inp[4]);
|
||||
dctx->s[2] = get_unaligned_le32(&inp[8]);
|
||||
dctx->s[3] = get_unaligned_le32(&inp[12]);
|
||||
inp += POLY1305_BLOCK_SIZE;
|
||||
len -= POLY1305_BLOCK_SIZE;
|
||||
acc += POLY1305_BLOCK_SIZE;
|
||||
dctx->sset = true;
|
||||
}
|
||||
blocks = srclen / (POLY1305_BLOCK_SIZE * 2);
|
||||
poly1305_2block_sse2(dctx->h.h, src, dctx->r.r, blocks,
|
||||
sctx->u);
|
||||
src += POLY1305_BLOCK_SIZE * 2 * blocks;
|
||||
srclen -= POLY1305_BLOCK_SIZE * 2 * blocks;
|
||||
}
|
||||
if (srclen >= POLY1305_BLOCK_SIZE) {
|
||||
poly1305_block_sse2(dctx->h.h, src, dctx->r.r, 1);
|
||||
srclen -= POLY1305_BLOCK_SIZE;
|
||||
}
|
||||
return srclen;
|
||||
return acc;
|
||||
}
|
||||
|
||||
static int poly1305_simd_update(struct shash_desc *desc,
|
||||
const u8 *src, unsigned int srclen)
|
||||
void poly1305_update_arch(struct poly1305_desc_ctx *dctx, const u8 *src,
|
||||
unsigned int srclen)
|
||||
{
|
||||
struct poly1305_desc_ctx *dctx = shash_desc_ctx(desc);
|
||||
unsigned int bytes;
|
||||
|
||||
/* kernel_fpu_begin/end is costly, use fallback for small updates */
|
||||
if (srclen <= 288 || !crypto_simd_usable())
|
||||
return crypto_poly1305_update(desc, src, srclen);
|
||||
|
||||
kernel_fpu_begin();
|
||||
unsigned int bytes, used;
|
||||
|
||||
if (unlikely(dctx->buflen)) {
|
||||
bytes = min(srclen, POLY1305_BLOCK_SIZE - dctx->buflen);
|
||||
|
|
@ -135,34 +179,76 @@ static int poly1305_simd_update(struct shash_desc *desc,
|
|||
dctx->buflen += bytes;
|
||||
|
||||
if (dctx->buflen == POLY1305_BLOCK_SIZE) {
|
||||
poly1305_simd_blocks(dctx, dctx->buf,
|
||||
POLY1305_BLOCK_SIZE);
|
||||
if (likely(!crypto_poly1305_setdctxkey(dctx, dctx->buf, POLY1305_BLOCK_SIZE)))
|
||||
poly1305_simd_blocks(&dctx->h, dctx->buf, POLY1305_BLOCK_SIZE, 1);
|
||||
dctx->buflen = 0;
|
||||
}
|
||||
}
|
||||
|
||||
if (likely(srclen >= POLY1305_BLOCK_SIZE)) {
|
||||
bytes = poly1305_simd_blocks(dctx, src, srclen);
|
||||
src += srclen - bytes;
|
||||
srclen = bytes;
|
||||
bytes = round_down(srclen, POLY1305_BLOCK_SIZE);
|
||||
srclen -= bytes;
|
||||
used = crypto_poly1305_setdctxkey(dctx, src, bytes);
|
||||
if (likely(bytes - used))
|
||||
poly1305_simd_blocks(&dctx->h, src + used, bytes - used, 1);
|
||||
src += bytes;
|
||||
}
|
||||
|
||||
kernel_fpu_end();
|
||||
|
||||
if (unlikely(srclen)) {
|
||||
dctx->buflen = srclen;
|
||||
memcpy(dctx->buf, src, srclen);
|
||||
}
|
||||
}
|
||||
EXPORT_SYMBOL(poly1305_update_arch);
|
||||
|
||||
void poly1305_final_arch(struct poly1305_desc_ctx *dctx, u8 *dst)
|
||||
{
|
||||
if (unlikely(dctx->buflen)) {
|
||||
dctx->buf[dctx->buflen++] = 1;
|
||||
memset(dctx->buf + dctx->buflen, 0,
|
||||
POLY1305_BLOCK_SIZE - dctx->buflen);
|
||||
poly1305_simd_blocks(&dctx->h, dctx->buf, POLY1305_BLOCK_SIZE, 0);
|
||||
}
|
||||
|
||||
poly1305_simd_emit(&dctx->h, dst, dctx->s);
|
||||
*dctx = (struct poly1305_desc_ctx){};
|
||||
}
|
||||
EXPORT_SYMBOL(poly1305_final_arch);
|
||||
|
||||
static int crypto_poly1305_init(struct shash_desc *desc)
|
||||
{
|
||||
struct poly1305_desc_ctx *dctx = shash_desc_ctx(desc);
|
||||
|
||||
*dctx = (struct poly1305_desc_ctx){};
|
||||
return 0;
|
||||
}
|
||||
|
||||
static int crypto_poly1305_update(struct shash_desc *desc,
|
||||
const u8 *src, unsigned int srclen)
|
||||
{
|
||||
struct poly1305_desc_ctx *dctx = shash_desc_ctx(desc);
|
||||
|
||||
poly1305_update_arch(dctx, src, srclen);
|
||||
return 0;
|
||||
}
|
||||
|
||||
static int crypto_poly1305_final(struct shash_desc *desc, u8 *dst)
|
||||
{
|
||||
struct poly1305_desc_ctx *dctx = shash_desc_ctx(desc);
|
||||
|
||||
if (unlikely(!dctx->sset))
|
||||
return -ENOKEY;
|
||||
|
||||
poly1305_final_arch(dctx, dst);
|
||||
return 0;
|
||||
}
|
||||
|
||||
static struct shash_alg alg = {
|
||||
.digestsize = POLY1305_DIGEST_SIZE,
|
||||
.init = poly1305_simd_init,
|
||||
.update = poly1305_simd_update,
|
||||
.init = crypto_poly1305_init,
|
||||
.update = crypto_poly1305_update,
|
||||
.final = crypto_poly1305_final,
|
||||
.descsize = sizeof(struct poly1305_simd_desc_ctx),
|
||||
.descsize = sizeof(struct poly1305_desc_ctx),
|
||||
.base = {
|
||||
.cra_name = "poly1305",
|
||||
.cra_driver_name = "poly1305-simd",
|
||||
|
|
@ -174,30 +260,33 @@ static struct shash_alg alg = {
|
|||
|
||||
static int __init poly1305_simd_mod_init(void)
|
||||
{
|
||||
if (!boot_cpu_has(X86_FEATURE_XMM2))
|
||||
return -ENODEV;
|
||||
|
||||
#ifdef CONFIG_AS_AVX2
|
||||
poly1305_use_avx2 = boot_cpu_has(X86_FEATURE_AVX) &&
|
||||
boot_cpu_has(X86_FEATURE_AVX2) &&
|
||||
cpu_has_xfeatures(XFEATURE_MASK_SSE | XFEATURE_MASK_YMM, NULL);
|
||||
alg.descsize = sizeof(struct poly1305_simd_desc_ctx);
|
||||
if (poly1305_use_avx2)
|
||||
alg.descsize += 10 * sizeof(u32);
|
||||
#endif
|
||||
return crypto_register_shash(&alg);
|
||||
if (IS_ENABLED(CONFIG_AS_AVX) && boot_cpu_has(X86_FEATURE_AVX) &&
|
||||
cpu_has_xfeatures(XFEATURE_MASK_SSE | XFEATURE_MASK_YMM, NULL))
|
||||
static_branch_enable(&poly1305_use_avx);
|
||||
if (IS_ENABLED(CONFIG_AS_AVX2) && boot_cpu_has(X86_FEATURE_AVX) &&
|
||||
boot_cpu_has(X86_FEATURE_AVX2) &&
|
||||
cpu_has_xfeatures(XFEATURE_MASK_SSE | XFEATURE_MASK_YMM, NULL))
|
||||
static_branch_enable(&poly1305_use_avx2);
|
||||
if (IS_ENABLED(CONFIG_AS_AVX512) && boot_cpu_has(X86_FEATURE_AVX) &&
|
||||
boot_cpu_has(X86_FEATURE_AVX2) && boot_cpu_has(X86_FEATURE_AVX512F) &&
|
||||
cpu_has_xfeatures(XFEATURE_MASK_SSE | XFEATURE_MASK_YMM | XFEATURE_MASK_AVX512, NULL) &&
|
||||
/* Skylake downclocks unacceptably much when using zmm, but later generations are fast. */
|
||||
boot_cpu_data.x86_model != INTEL_FAM6_SKYLAKE_X)
|
||||
static_branch_enable(&poly1305_use_avx512);
|
||||
return IS_REACHABLE(CONFIG_CRYPTO_HASH) ? crypto_register_shash(&alg) : 0;
|
||||
}
|
||||
|
||||
static void __exit poly1305_simd_mod_exit(void)
|
||||
{
|
||||
crypto_unregister_shash(&alg);
|
||||
if (IS_REACHABLE(CONFIG_CRYPTO_HASH))
|
||||
crypto_unregister_shash(&alg);
|
||||
}
|
||||
|
||||
module_init(poly1305_simd_mod_init);
|
||||
module_exit(poly1305_simd_mod_exit);
|
||||
|
||||
MODULE_LICENSE("GPL");
|
||||
MODULE_AUTHOR("Martin Willi <martin@strongswan.org>");
|
||||
MODULE_AUTHOR("Jason A. Donenfeld <Jason@zx2c4.com>");
|
||||
MODULE_DESCRIPTION("Poly1305 authenticator");
|
||||
MODULE_ALIAS_CRYPTO("poly1305");
|
||||
MODULE_ALIAS_CRYPTO("poly1305-simd");
|
||||
|
|
|
|||
|
|
@ -136,8 +136,6 @@ config CRYPTO_USER
|
|||
Userspace configuration for cryptographic instantiations such as
|
||||
cbc(aes).
|
||||
|
||||
if CRYPTO_MANAGER2
|
||||
|
||||
config CRYPTO_MANAGER_DISABLE_TESTS
|
||||
bool "Disable run-time self tests"
|
||||
default y
|
||||
|
|
@ -147,7 +145,7 @@ config CRYPTO_MANAGER_DISABLE_TESTS
|
|||
|
||||
config CRYPTO_MANAGER_EXTRA_TESTS
|
||||
bool "Enable extra run-time crypto self tests"
|
||||
depends on DEBUG_KERNEL && !CRYPTO_MANAGER_DISABLE_TESTS
|
||||
depends on DEBUG_KERNEL && !CRYPTO_MANAGER_DISABLE_TESTS && CRYPTO_MANAGER
|
||||
help
|
||||
Enable extra run-time self tests of registered crypto algorithms,
|
||||
including randomized fuzz tests.
|
||||
|
|
@ -155,8 +153,6 @@ config CRYPTO_MANAGER_EXTRA_TESTS
|
|||
This is intended for developer use only, as these tests take much
|
||||
longer to run than the normal self tests.
|
||||
|
||||
endif # if CRYPTO_MANAGER2
|
||||
|
||||
config CRYPTO_GF128MUL
|
||||
tristate
|
||||
|
||||
|
|
@ -264,6 +260,17 @@ config CRYPTO_ECRDSA
|
|||
standard algorithms (called GOST algorithms). Only signature verification
|
||||
is implemented.
|
||||
|
||||
config CRYPTO_CURVE25519
|
||||
tristate "Curve25519 algorithm"
|
||||
select CRYPTO_KPP
|
||||
select CRYPTO_LIB_CURVE25519_GENERIC
|
||||
|
||||
config CRYPTO_CURVE25519_X86
|
||||
tristate "x86_64 accelerated Curve25519 scalar multiplication library"
|
||||
depends on X86 && 64BIT
|
||||
select CRYPTO_LIB_CURVE25519_GENERIC
|
||||
select CRYPTO_ARCH_HAVE_LIB_CURVE25519
|
||||
|
||||
comment "Authenticated Encryption with Associated Data"
|
||||
|
||||
config CRYPTO_CCM
|
||||
|
|
@ -446,7 +453,7 @@ config CRYPTO_KEYWRAP
|
|||
config CRYPTO_NHPOLY1305
|
||||
tristate
|
||||
select CRYPTO_HASH
|
||||
select CRYPTO_POLY1305
|
||||
select CRYPTO_LIB_POLY1305_GENERIC
|
||||
|
||||
config CRYPTO_NHPOLY1305_SSE2
|
||||
tristate "NHPoly1305 hash function (x86_64 SSE2 implementation)"
|
||||
|
|
@ -467,7 +474,7 @@ config CRYPTO_NHPOLY1305_AVX2
|
|||
config CRYPTO_ADIANTUM
|
||||
tristate "Adiantum support"
|
||||
select CRYPTO_CHACHA20
|
||||
select CRYPTO_POLY1305
|
||||
select CRYPTO_LIB_POLY1305_GENERIC
|
||||
select CRYPTO_NHPOLY1305
|
||||
select CRYPTO_MANAGER
|
||||
help
|
||||
|
|
@ -727,6 +734,7 @@ config CRYPTO_GHASH
|
|||
config CRYPTO_POLY1305
|
||||
tristate "Poly1305 authenticator algorithm"
|
||||
select CRYPTO_HASH
|
||||
select CRYPTO_LIB_POLY1305_GENERIC
|
||||
help
|
||||
Poly1305 authenticator algorithm, RFC7539.
|
||||
|
||||
|
|
@ -737,7 +745,8 @@ config CRYPTO_POLY1305
|
|||
config CRYPTO_POLY1305_X86_64
|
||||
tristate "Poly1305 authenticator algorithm (x86_64/SSE2/AVX2)"
|
||||
depends on X86 && 64BIT
|
||||
select CRYPTO_POLY1305
|
||||
select CRYPTO_LIB_POLY1305_GENERIC
|
||||
select CRYPTO_ARCH_HAVE_LIB_POLY1305
|
||||
help
|
||||
Poly1305 authenticator algorithm, RFC7539.
|
||||
|
||||
|
|
@ -746,6 +755,11 @@ config CRYPTO_POLY1305_X86_64
|
|||
in IETF protocols. This is the x86_64 assembler implementation using SIMD
|
||||
instructions.
|
||||
|
||||
config CRYPTO_POLY1305_MIPS
|
||||
tristate "Poly1305 authenticator algorithm (MIPS optimized)"
|
||||
depends on MIPS
|
||||
select CRYPTO_ARCH_HAVE_LIB_POLY1305
|
||||
|
||||
config CRYPTO_MD4
|
||||
tristate "MD4 digest algorithm"
|
||||
select CRYPTO_HASH
|
||||
|
|
@ -1434,6 +1448,7 @@ config CRYPTO_SALSA20
|
|||
|
||||
config CRYPTO_CHACHA20
|
||||
tristate "ChaCha stream cipher algorithms"
|
||||
select CRYPTO_LIB_CHACHA_GENERIC
|
||||
select CRYPTO_BLKCIPHER
|
||||
help
|
||||
The ChaCha20, XChaCha20, and XChaCha12 stream cipher algorithms.
|
||||
|
|
@ -1457,11 +1472,18 @@ config CRYPTO_CHACHA20_X86_64
|
|||
tristate "ChaCha stream cipher algorithms (x86_64/SSSE3/AVX2/AVX-512VL)"
|
||||
depends on X86 && 64BIT
|
||||
select CRYPTO_BLKCIPHER
|
||||
select CRYPTO_CHACHA20
|
||||
select CRYPTO_LIB_CHACHA_GENERIC
|
||||
select CRYPTO_ARCH_HAVE_LIB_CHACHA
|
||||
help
|
||||
SSSE3, AVX2, and AVX-512VL optimized implementations of the ChaCha20,
|
||||
XChaCha20, and XChaCha12 stream ciphers.
|
||||
|
||||
config CRYPTO_CHACHA_MIPS
|
||||
tristate "ChaCha stream cipher algorithms (MIPS 32r2 optimized)"
|
||||
depends on CPU_MIPS32_R2
|
||||
select CRYPTO_BLKCIPHER
|
||||
select CRYPTO_ARCH_HAVE_LIB_CHACHA
|
||||
|
||||
config CRYPTO_SEED
|
||||
tristate "SEED cipher algorithm"
|
||||
select CRYPTO_ALGAPI
|
||||
|
|
|
|||
|
|
@ -168,6 +168,7 @@ obj-$(CONFIG_CRYPTO_ZSTD) += zstd.o
|
|||
obj-$(CONFIG_CRYPTO_OFB) += ofb.o
|
||||
obj-$(CONFIG_CRYPTO_ECC) += ecc.o
|
||||
obj-$(CONFIG_CRYPTO_ESSIV) += essiv.o
|
||||
obj-$(CONFIG_CRYPTO_CURVE25519) += curve25519-generic.o
|
||||
|
||||
ecdh_generic-y += ecdh.o
|
||||
ecdh_generic-y += ecdh_helper.o
|
||||
|
|
|
|||
|
|
@ -33,6 +33,7 @@
|
|||
#include <crypto/b128ops.h>
|
||||
#include <crypto/chacha.h>
|
||||
#include <crypto/internal/hash.h>
|
||||
#include <crypto/internal/poly1305.h>
|
||||
#include <crypto/internal/skcipher.h>
|
||||
#include <crypto/nhpoly1305.h>
|
||||
#include <crypto/scatterwalk.h>
|
||||
|
|
@ -71,7 +72,7 @@ struct adiantum_tfm_ctx {
|
|||
struct crypto_skcipher *streamcipher;
|
||||
struct crypto_cipher *blockcipher;
|
||||
struct crypto_shash *hash;
|
||||
struct poly1305_key header_hash_key;
|
||||
struct poly1305_core_key header_hash_key;
|
||||
};
|
||||
|
||||
struct adiantum_request_ctx {
|
||||
|
|
@ -242,13 +243,13 @@ static void adiantum_hash_header(struct skcipher_request *req)
|
|||
|
||||
BUILD_BUG_ON(sizeof(header) % POLY1305_BLOCK_SIZE != 0);
|
||||
poly1305_core_blocks(&state, &tctx->header_hash_key,
|
||||
&header, sizeof(header) / POLY1305_BLOCK_SIZE);
|
||||
&header, sizeof(header) / POLY1305_BLOCK_SIZE, 1);
|
||||
|
||||
BUILD_BUG_ON(TWEAK_SIZE % POLY1305_BLOCK_SIZE != 0);
|
||||
poly1305_core_blocks(&state, &tctx->header_hash_key, req->iv,
|
||||
TWEAK_SIZE / POLY1305_BLOCK_SIZE);
|
||||
TWEAK_SIZE / POLY1305_BLOCK_SIZE, 1);
|
||||
|
||||
poly1305_core_emit(&state, &rctx->header_hash);
|
||||
poly1305_core_emit(&state, NULL, &rctx->header_hash);
|
||||
}
|
||||
|
||||
/* Hash the left-hand part (the "bulk") of the message using NHPoly1305 */
|
||||
|
|
|
|||
|
|
@ -8,29 +8,10 @@
|
|||
|
||||
#include <asm/unaligned.h>
|
||||
#include <crypto/algapi.h>
|
||||
#include <crypto/chacha.h>
|
||||
#include <crypto/internal/chacha.h>
|
||||
#include <crypto/internal/skcipher.h>
|
||||
#include <linux/module.h>
|
||||
|
||||
static void chacha_docrypt(u32 *state, u8 *dst, const u8 *src,
|
||||
unsigned int bytes, int nrounds)
|
||||
{
|
||||
/* aligned to potentially speed up crypto_xor() */
|
||||
u8 stream[CHACHA_BLOCK_SIZE] __aligned(sizeof(long));
|
||||
|
||||
while (bytes >= CHACHA_BLOCK_SIZE) {
|
||||
chacha_block(state, stream, nrounds);
|
||||
crypto_xor_cpy(dst, src, stream, CHACHA_BLOCK_SIZE);
|
||||
bytes -= CHACHA_BLOCK_SIZE;
|
||||
dst += CHACHA_BLOCK_SIZE;
|
||||
src += CHACHA_BLOCK_SIZE;
|
||||
}
|
||||
if (bytes) {
|
||||
chacha_block(state, stream, nrounds);
|
||||
crypto_xor_cpy(dst, src, stream, bytes);
|
||||
}
|
||||
}
|
||||
|
||||
static int chacha_stream_xor(struct skcipher_request *req,
|
||||
const struct chacha_ctx *ctx, const u8 *iv)
|
||||
{
|
||||
|
|
@ -40,7 +21,7 @@ static int chacha_stream_xor(struct skcipher_request *req,
|
|||
|
||||
err = skcipher_walk_virt(&walk, req, false);
|
||||
|
||||
crypto_chacha_init(state, ctx, iv);
|
||||
chacha_init_generic(state, ctx->key, iv);
|
||||
|
||||
while (walk.nbytes > 0) {
|
||||
unsigned int nbytes = walk.nbytes;
|
||||
|
|
@ -48,75 +29,23 @@ static int chacha_stream_xor(struct skcipher_request *req,
|
|||
if (nbytes < walk.total)
|
||||
nbytes = round_down(nbytes, CHACHA_BLOCK_SIZE);
|
||||
|
||||
chacha_docrypt(state, walk.dst.virt.addr, walk.src.virt.addr,
|
||||
nbytes, ctx->nrounds);
|
||||
chacha_crypt_generic(state, walk.dst.virt.addr,
|
||||
walk.src.virt.addr, nbytes, ctx->nrounds);
|
||||
err = skcipher_walk_done(&walk, walk.nbytes - nbytes);
|
||||
}
|
||||
|
||||
return err;
|
||||
}
|
||||
|
||||
void crypto_chacha_init(u32 *state, const struct chacha_ctx *ctx, const u8 *iv)
|
||||
{
|
||||
state[0] = 0x61707865; /* "expa" */
|
||||
state[1] = 0x3320646e; /* "nd 3" */
|
||||
state[2] = 0x79622d32; /* "2-by" */
|
||||
state[3] = 0x6b206574; /* "te k" */
|
||||
state[4] = ctx->key[0];
|
||||
state[5] = ctx->key[1];
|
||||
state[6] = ctx->key[2];
|
||||
state[7] = ctx->key[3];
|
||||
state[8] = ctx->key[4];
|
||||
state[9] = ctx->key[5];
|
||||
state[10] = ctx->key[6];
|
||||
state[11] = ctx->key[7];
|
||||
state[12] = get_unaligned_le32(iv + 0);
|
||||
state[13] = get_unaligned_le32(iv + 4);
|
||||
state[14] = get_unaligned_le32(iv + 8);
|
||||
state[15] = get_unaligned_le32(iv + 12);
|
||||
}
|
||||
EXPORT_SYMBOL_GPL(crypto_chacha_init);
|
||||
|
||||
static int chacha_setkey(struct crypto_skcipher *tfm, const u8 *key,
|
||||
unsigned int keysize, int nrounds)
|
||||
{
|
||||
struct chacha_ctx *ctx = crypto_skcipher_ctx(tfm);
|
||||
int i;
|
||||
|
||||
if (keysize != CHACHA_KEY_SIZE)
|
||||
return -EINVAL;
|
||||
|
||||
for (i = 0; i < ARRAY_SIZE(ctx->key); i++)
|
||||
ctx->key[i] = get_unaligned_le32(key + i * sizeof(u32));
|
||||
|
||||
ctx->nrounds = nrounds;
|
||||
return 0;
|
||||
}
|
||||
|
||||
int crypto_chacha20_setkey(struct crypto_skcipher *tfm, const u8 *key,
|
||||
unsigned int keysize)
|
||||
{
|
||||
return chacha_setkey(tfm, key, keysize, 20);
|
||||
}
|
||||
EXPORT_SYMBOL_GPL(crypto_chacha20_setkey);
|
||||
|
||||
int crypto_chacha12_setkey(struct crypto_skcipher *tfm, const u8 *key,
|
||||
unsigned int keysize)
|
||||
{
|
||||
return chacha_setkey(tfm, key, keysize, 12);
|
||||
}
|
||||
EXPORT_SYMBOL_GPL(crypto_chacha12_setkey);
|
||||
|
||||
int crypto_chacha_crypt(struct skcipher_request *req)
|
||||
static int crypto_chacha_crypt(struct skcipher_request *req)
|
||||
{
|
||||
struct crypto_skcipher *tfm = crypto_skcipher_reqtfm(req);
|
||||
struct chacha_ctx *ctx = crypto_skcipher_ctx(tfm);
|
||||
|
||||
return chacha_stream_xor(req, ctx, req->iv);
|
||||
}
|
||||
EXPORT_SYMBOL_GPL(crypto_chacha_crypt);
|
||||
|
||||
int crypto_xchacha_crypt(struct skcipher_request *req)
|
||||
static int crypto_xchacha_crypt(struct skcipher_request *req)
|
||||
{
|
||||
struct crypto_skcipher *tfm = crypto_skcipher_reqtfm(req);
|
||||
struct chacha_ctx *ctx = crypto_skcipher_ctx(tfm);
|
||||
|
|
@ -125,8 +54,8 @@ int crypto_xchacha_crypt(struct skcipher_request *req)
|
|||
u8 real_iv[16];
|
||||
|
||||
/* Compute the subkey given the original key and first 128 nonce bits */
|
||||
crypto_chacha_init(state, ctx, req->iv);
|
||||
hchacha_block(state, subctx.key, ctx->nrounds);
|
||||
chacha_init_generic(state, ctx->key, req->iv);
|
||||
hchacha_block_generic(state, subctx.key, ctx->nrounds);
|
||||
subctx.nrounds = ctx->nrounds;
|
||||
|
||||
/* Build the real IV */
|
||||
|
|
@ -136,7 +65,6 @@ int crypto_xchacha_crypt(struct skcipher_request *req)
|
|||
/* Generate the stream and XOR it with the data */
|
||||
return chacha_stream_xor(req, &subctx, real_iv);
|
||||
}
|
||||
EXPORT_SYMBOL_GPL(crypto_xchacha_crypt);
|
||||
|
||||
static struct skcipher_alg algs[] = {
|
||||
{
|
||||
|
|
@ -151,7 +79,7 @@ static struct skcipher_alg algs[] = {
|
|||
.max_keysize = CHACHA_KEY_SIZE,
|
||||
.ivsize = CHACHA_IV_SIZE,
|
||||
.chunksize = CHACHA_BLOCK_SIZE,
|
||||
.setkey = crypto_chacha20_setkey,
|
||||
.setkey = chacha20_setkey,
|
||||
.encrypt = crypto_chacha_crypt,
|
||||
.decrypt = crypto_chacha_crypt,
|
||||
}, {
|
||||
|
|
@ -166,7 +94,7 @@ static struct skcipher_alg algs[] = {
|
|||
.max_keysize = CHACHA_KEY_SIZE,
|
||||
.ivsize = XCHACHA_IV_SIZE,
|
||||
.chunksize = CHACHA_BLOCK_SIZE,
|
||||
.setkey = crypto_chacha20_setkey,
|
||||
.setkey = chacha20_setkey,
|
||||
.encrypt = crypto_xchacha_crypt,
|
||||
.decrypt = crypto_xchacha_crypt,
|
||||
}, {
|
||||
|
|
@ -181,7 +109,7 @@ static struct skcipher_alg algs[] = {
|
|||
.max_keysize = CHACHA_KEY_SIZE,
|
||||
.ivsize = XCHACHA_IV_SIZE,
|
||||
.chunksize = CHACHA_BLOCK_SIZE,
|
||||
.setkey = crypto_chacha12_setkey,
|
||||
.setkey = chacha12_setkey,
|
||||
.encrypt = crypto_xchacha_crypt,
|
||||
.decrypt = crypto_xchacha_crypt,
|
||||
}
|
||||
|
|
|
|||
90
crypto/curve25519-generic.c
Normal file
90
crypto/curve25519-generic.c
Normal file
|
|
@ -0,0 +1,90 @@
|
|||
// SPDX-License-Identifier: GPL-2.0-or-later
|
||||
|
||||
#include <crypto/curve25519.h>
|
||||
#include <crypto/internal/kpp.h>
|
||||
#include <crypto/kpp.h>
|
||||
#include <linux/module.h>
|
||||
#include <linux/scatterlist.h>
|
||||
|
||||
static int curve25519_set_secret(struct crypto_kpp *tfm, const void *buf,
|
||||
unsigned int len)
|
||||
{
|
||||
u8 *secret = kpp_tfm_ctx(tfm);
|
||||
|
||||
if (!len)
|
||||
curve25519_generate_secret(secret);
|
||||
else if (len == CURVE25519_KEY_SIZE &&
|
||||
crypto_memneq(buf, curve25519_null_point, CURVE25519_KEY_SIZE))
|
||||
memcpy(secret, buf, CURVE25519_KEY_SIZE);
|
||||
else
|
||||
return -EINVAL;
|
||||
return 0;
|
||||
}
|
||||
|
||||
static int curve25519_compute_value(struct kpp_request *req)
|
||||
{
|
||||
struct crypto_kpp *tfm = crypto_kpp_reqtfm(req);
|
||||
const u8 *secret = kpp_tfm_ctx(tfm);
|
||||
u8 public_key[CURVE25519_KEY_SIZE];
|
||||
u8 buf[CURVE25519_KEY_SIZE];
|
||||
int copied, nbytes;
|
||||
u8 const *bp;
|
||||
|
||||
if (req->src) {
|
||||
copied = sg_copy_to_buffer(req->src,
|
||||
sg_nents_for_len(req->src,
|
||||
CURVE25519_KEY_SIZE),
|
||||
public_key, CURVE25519_KEY_SIZE);
|
||||
if (copied != CURVE25519_KEY_SIZE)
|
||||
return -EINVAL;
|
||||
bp = public_key;
|
||||
} else {
|
||||
bp = curve25519_base_point;
|
||||
}
|
||||
|
||||
curve25519_generic(buf, secret, bp);
|
||||
|
||||
/* might want less than we've got */
|
||||
nbytes = min_t(size_t, CURVE25519_KEY_SIZE, req->dst_len);
|
||||
copied = sg_copy_from_buffer(req->dst, sg_nents_for_len(req->dst,
|
||||
nbytes),
|
||||
buf, nbytes);
|
||||
if (copied != nbytes)
|
||||
return -EINVAL;
|
||||
return 0;
|
||||
}
|
||||
|
||||
static unsigned int curve25519_max_size(struct crypto_kpp *tfm)
|
||||
{
|
||||
return CURVE25519_KEY_SIZE;
|
||||
}
|
||||
|
||||
static struct kpp_alg curve25519_alg = {
|
||||
.base.cra_name = "curve25519",
|
||||
.base.cra_driver_name = "curve25519-generic",
|
||||
.base.cra_priority = 100,
|
||||
.base.cra_module = THIS_MODULE,
|
||||
.base.cra_ctxsize = CURVE25519_KEY_SIZE,
|
||||
|
||||
.set_secret = curve25519_set_secret,
|
||||
.generate_public_key = curve25519_compute_value,
|
||||
.compute_shared_secret = curve25519_compute_value,
|
||||
.max_size = curve25519_max_size,
|
||||
};
|
||||
|
||||
static int curve25519_init(void)
|
||||
{
|
||||
return crypto_register_kpp(&curve25519_alg);
|
||||
}
|
||||
|
||||
static void curve25519_exit(void)
|
||||
{
|
||||
crypto_unregister_kpp(&curve25519_alg);
|
||||
}
|
||||
|
||||
subsys_initcall(curve25519_init);
|
||||
module_exit(curve25519_exit);
|
||||
|
||||
MODULE_ALIAS_CRYPTO("curve25519");
|
||||
MODULE_ALIAS_CRYPTO("curve25519-generic");
|
||||
MODULE_LICENSE("GPL");
|
||||
|
|
@ -44,13 +44,7 @@
|
|||
#include <linux/crypto.h>
|
||||
#include <crypto/internal/rng.h>
|
||||
|
||||
struct rand_data;
|
||||
int jent_read_entropy(struct rand_data *ec, unsigned char *data,
|
||||
unsigned int len);
|
||||
int jent_entropy_init(void);
|
||||
struct rand_data *jent_entropy_collector_alloc(unsigned int osr,
|
||||
unsigned int flags);
|
||||
void jent_entropy_collector_free(struct rand_data *entropy_collector);
|
||||
#include "jitterentropy.h"
|
||||
|
||||
/***************************************************************************
|
||||
* Helper function
|
||||
|
|
@ -114,6 +108,7 @@ void jent_get_nstime(__u64 *out)
|
|||
struct jitterentropy {
|
||||
spinlock_t jent_lock;
|
||||
struct rand_data *entropy_collector;
|
||||
unsigned int reset_cnt;
|
||||
};
|
||||
|
||||
static int jent_kcapi_init(struct crypto_tfm *tfm)
|
||||
|
|
@ -148,7 +143,33 @@ static int jent_kcapi_random(struct crypto_rng *tfm,
|
|||
int ret = 0;
|
||||
|
||||
spin_lock(&rng->jent_lock);
|
||||
|
||||
/* Return a permanent error in case we had too many resets in a row. */
|
||||
if (rng->reset_cnt > (1<<10)) {
|
||||
ret = -EFAULT;
|
||||
goto out;
|
||||
}
|
||||
|
||||
ret = jent_read_entropy(rng->entropy_collector, rdata, dlen);
|
||||
|
||||
/* Reset RNG in case of health failures */
|
||||
if (ret < -1) {
|
||||
pr_warn_ratelimited("Reset Jitter RNG due to health test failure: %s failure\n",
|
||||
(ret == -2) ? "Repetition Count Test" :
|
||||
"Adaptive Proportion Test");
|
||||
|
||||
rng->reset_cnt++;
|
||||
|
||||
ret = -EAGAIN;
|
||||
} else {
|
||||
rng->reset_cnt = 0;
|
||||
|
||||
/* Convert the Jitter RNG error into a usable error code */
|
||||
if (ret == -1)
|
||||
ret = -EINVAL;
|
||||
}
|
||||
|
||||
out:
|
||||
spin_unlock(&rng->jent_lock);
|
||||
|
||||
return ret;
|
||||
|
|
|
|||
|
|
@ -2,7 +2,7 @@
|
|||
* Non-physical true random number generator based on timing jitter --
|
||||
* Jitter RNG standalone code.
|
||||
*
|
||||
* Copyright Stephan Mueller <smueller@chronox.de>, 2015 - 2019
|
||||
* Copyright Stephan Mueller <smueller@chronox.de>, 2015 - 2020
|
||||
*
|
||||
* Design
|
||||
* ======
|
||||
|
|
@ -47,7 +47,7 @@
|
|||
|
||||
/*
|
||||
* This Jitterentropy RNG is based on the jitterentropy library
|
||||
* version 2.1.2 provided at http://www.chronox.de/jent.html
|
||||
* version 2.2.0 provided at http://www.chronox.de/jent.html
|
||||
*/
|
||||
|
||||
#ifdef __OPTIMIZE__
|
||||
|
|
@ -83,6 +83,22 @@ struct rand_data {
|
|||
unsigned int memblocksize; /* Size of one memory block in bytes */
|
||||
unsigned int memaccessloops; /* Number of memory accesses per random
|
||||
* bit generation */
|
||||
|
||||
/* Repetition Count Test */
|
||||
int rct_count; /* Number of stuck values */
|
||||
|
||||
/* Adaptive Proportion Test for a significance level of 2^-30 */
|
||||
#define JENT_APT_CUTOFF 325 /* Taken from SP800-90B sec 4.4.2 */
|
||||
#define JENT_APT_WINDOW_SIZE 512 /* Data window size */
|
||||
/* LSB of time stamp to process */
|
||||
#define JENT_APT_LSB 16
|
||||
#define JENT_APT_WORD_MASK (JENT_APT_LSB - 1)
|
||||
unsigned int apt_observations; /* Number of collected observations */
|
||||
unsigned int apt_count; /* APT counter */
|
||||
unsigned int apt_base; /* APT base reference */
|
||||
unsigned int apt_base_set:1; /* APT base reference set? */
|
||||
|
||||
unsigned int health_failure:1; /* Permanent health failure */
|
||||
};
|
||||
|
||||
/* Flags that can be used to initialize the RNG */
|
||||
|
|
@ -98,17 +114,201 @@ struct rand_data {
|
|||
* variations (2nd derivation of time is
|
||||
* zero). */
|
||||
#define JENT_ESTUCK 8 /* Too many stuck results during init. */
|
||||
#define JENT_EHEALTH 9 /* Health test failed during initialization */
|
||||
#define JENT_ERCT 10 /* RCT failed during initialization */
|
||||
|
||||
#include "jitterentropy.h"
|
||||
|
||||
/***************************************************************************
|
||||
* Helper functions
|
||||
* Adaptive Proportion Test
|
||||
*
|
||||
* This test complies with SP800-90B section 4.4.2.
|
||||
***************************************************************************/
|
||||
|
||||
void jent_get_nstime(__u64 *out);
|
||||
void *jent_zalloc(unsigned int len);
|
||||
void jent_zfree(void *ptr);
|
||||
int jent_fips_enabled(void);
|
||||
void jent_panic(char *s);
|
||||
void jent_memcpy(void *dest, const void *src, unsigned int n);
|
||||
/**
|
||||
* Reset the APT counter
|
||||
*
|
||||
* @ec [in] Reference to entropy collector
|
||||
*/
|
||||
static void jent_apt_reset(struct rand_data *ec, unsigned int delta_masked)
|
||||
{
|
||||
/* Reset APT counter */
|
||||
ec->apt_count = 0;
|
||||
ec->apt_base = delta_masked;
|
||||
ec->apt_observations = 0;
|
||||
}
|
||||
|
||||
/**
|
||||
* Insert a new entropy event into APT
|
||||
*
|
||||
* @ec [in] Reference to entropy collector
|
||||
* @delta_masked [in] Masked time delta to process
|
||||
*/
|
||||
static void jent_apt_insert(struct rand_data *ec, unsigned int delta_masked)
|
||||
{
|
||||
/* Initialize the base reference */
|
||||
if (!ec->apt_base_set) {
|
||||
ec->apt_base = delta_masked;
|
||||
ec->apt_base_set = 1;
|
||||
return;
|
||||
}
|
||||
|
||||
if (delta_masked == ec->apt_base) {
|
||||
ec->apt_count++;
|
||||
|
||||
if (ec->apt_count >= JENT_APT_CUTOFF)
|
||||
ec->health_failure = 1;
|
||||
}
|
||||
|
||||
ec->apt_observations++;
|
||||
|
||||
if (ec->apt_observations >= JENT_APT_WINDOW_SIZE)
|
||||
jent_apt_reset(ec, delta_masked);
|
||||
}
|
||||
|
||||
/***************************************************************************
|
||||
* Stuck Test and its use as Repetition Count Test
|
||||
*
|
||||
* The Jitter RNG uses an enhanced version of the Repetition Count Test
|
||||
* (RCT) specified in SP800-90B section 4.4.1. Instead of counting identical
|
||||
* back-to-back values, the input to the RCT is the counting of the stuck
|
||||
* values during the generation of one Jitter RNG output block.
|
||||
*
|
||||
* The RCT is applied with an alpha of 2^{-30} compliant to FIPS 140-2 IG 9.8.
|
||||
*
|
||||
* During the counting operation, the Jitter RNG always calculates the RCT
|
||||
* cut-off value of C. If that value exceeds the allowed cut-off value,
|
||||
* the Jitter RNG output block will be calculated completely but discarded at
|
||||
* the end. The caller of the Jitter RNG is informed with an error code.
|
||||
***************************************************************************/
|
||||
|
||||
/**
|
||||
* Repetition Count Test as defined in SP800-90B section 4.4.1
|
||||
*
|
||||
* @ec [in] Reference to entropy collector
|
||||
* @stuck [in] Indicator whether the value is stuck
|
||||
*/
|
||||
static void jent_rct_insert(struct rand_data *ec, int stuck)
|
||||
{
|
||||
/*
|
||||
* If we have a count less than zero, a previous RCT round identified
|
||||
* a failure. We will not overwrite it.
|
||||
*/
|
||||
if (ec->rct_count < 0)
|
||||
return;
|
||||
|
||||
if (stuck) {
|
||||
ec->rct_count++;
|
||||
|
||||
/*
|
||||
* The cutoff value is based on the following consideration:
|
||||
* alpha = 2^-30 as recommended in FIPS 140-2 IG 9.8.
|
||||
* In addition, we require an entropy value H of 1/OSR as this
|
||||
* is the minimum entropy required to provide full entropy.
|
||||
* Note, we collect 64 * OSR deltas for inserting them into
|
||||
* the entropy pool which should then have (close to) 64 bits
|
||||
* of entropy.
|
||||
*
|
||||
* Note, ec->rct_count (which equals to value B in the pseudo
|
||||
* code of SP800-90B section 4.4.1) starts with zero. Hence
|
||||
* we need to subtract one from the cutoff value as calculated
|
||||
* following SP800-90B.
|
||||
*/
|
||||
if ((unsigned int)ec->rct_count >= (31 * ec->osr)) {
|
||||
ec->rct_count = -1;
|
||||
ec->health_failure = 1;
|
||||
}
|
||||
} else {
|
||||
ec->rct_count = 0;
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Is there an RCT health test failure?
|
||||
*
|
||||
* @ec [in] Reference to entropy collector
|
||||
*
|
||||
* @return
|
||||
* 0 No health test failure
|
||||
* 1 Permanent health test failure
|
||||
*/
|
||||
static int jent_rct_failure(struct rand_data *ec)
|
||||
{
|
||||
if (ec->rct_count < 0)
|
||||
return 1;
|
||||
return 0;
|
||||
}
|
||||
|
||||
static inline __u64 jent_delta(__u64 prev, __u64 next)
|
||||
{
|
||||
#define JENT_UINT64_MAX (__u64)(~((__u64) 0))
|
||||
return (prev < next) ? (next - prev) :
|
||||
(JENT_UINT64_MAX - prev + 1 + next);
|
||||
}
|
||||
|
||||
/**
|
||||
* Stuck test by checking the:
|
||||
* 1st derivative of the jitter measurement (time delta)
|
||||
* 2nd derivative of the jitter measurement (delta of time deltas)
|
||||
* 3rd derivative of the jitter measurement (delta of delta of time deltas)
|
||||
*
|
||||
* All values must always be non-zero.
|
||||
*
|
||||
* @ec [in] Reference to entropy collector
|
||||
* @current_delta [in] Jitter time delta
|
||||
*
|
||||
* @return
|
||||
* 0 jitter measurement not stuck (good bit)
|
||||
* 1 jitter measurement stuck (reject bit)
|
||||
*/
|
||||
static int jent_stuck(struct rand_data *ec, __u64 current_delta)
|
||||
{
|
||||
__u64 delta2 = jent_delta(ec->last_delta, current_delta);
|
||||
__u64 delta3 = jent_delta(ec->last_delta2, delta2);
|
||||
unsigned int delta_masked = current_delta & JENT_APT_WORD_MASK;
|
||||
|
||||
ec->last_delta = current_delta;
|
||||
ec->last_delta2 = delta2;
|
||||
|
||||
/*
|
||||
* Insert the result of the comparison of two back-to-back time
|
||||
* deltas.
|
||||
*/
|
||||
jent_apt_insert(ec, delta_masked);
|
||||
|
||||
if (!current_delta || !delta2 || !delta3) {
|
||||
/* RCT with a stuck bit */
|
||||
jent_rct_insert(ec, 1);
|
||||
return 1;
|
||||
}
|
||||
|
||||
/* RCT with a non-stuck bit */
|
||||
jent_rct_insert(ec, 0);
|
||||
|
||||
return 0;
|
||||
}
|
||||
|
||||
/**
|
||||
* Report any health test failures
|
||||
*
|
||||
* @ec [in] Reference to entropy collector
|
||||
*
|
||||
* @return
|
||||
* 0 No health test failure
|
||||
* 1 Permanent health test failure
|
||||
*/
|
||||
static int jent_health_failure(struct rand_data *ec)
|
||||
{
|
||||
/* Test is only enabled in FIPS mode */
|
||||
if (!jent_fips_enabled())
|
||||
return 0;
|
||||
|
||||
return ec->health_failure;
|
||||
}
|
||||
|
||||
/***************************************************************************
|
||||
* Noise sources
|
||||
***************************************************************************/
|
||||
|
||||
/**
|
||||
* Update of the loop count used for the next round of
|
||||
|
|
@ -153,10 +353,6 @@ static __u64 jent_loop_shuffle(struct rand_data *ec,
|
|||
return (shuffle + (1<<min));
|
||||
}
|
||||
|
||||
/***************************************************************************
|
||||
* Noise sources
|
||||
***************************************************************************/
|
||||
|
||||
/**
|
||||
* CPU Jitter noise source -- this is the noise source based on the CPU
|
||||
* execution time jitter
|
||||
|
|
@ -171,18 +367,19 @@ static __u64 jent_loop_shuffle(struct rand_data *ec,
|
|||
* the CPU execution time jitter. Any change to the loop in this function
|
||||
* implies that careful retesting must be done.
|
||||
*
|
||||
* Input:
|
||||
* @ec entropy collector struct -- may be NULL
|
||||
* @time time stamp to be injected
|
||||
* @loop_cnt if a value not equal to 0 is set, use the given value as number of
|
||||
* loops to perform the folding
|
||||
* @ec [in] entropy collector struct
|
||||
* @time [in] time stamp to be injected
|
||||
* @loop_cnt [in] if a value not equal to 0 is set, use the given value as
|
||||
* number of loops to perform the folding
|
||||
* @stuck [in] Is the time stamp identified as stuck?
|
||||
*
|
||||
* Output:
|
||||
* updated ec->data
|
||||
*
|
||||
* @return Number of loops the folding operation is performed
|
||||
*/
|
||||
static __u64 jent_lfsr_time(struct rand_data *ec, __u64 time, __u64 loop_cnt)
|
||||
static void jent_lfsr_time(struct rand_data *ec, __u64 time, __u64 loop_cnt,
|
||||
int stuck)
|
||||
{
|
||||
unsigned int i;
|
||||
__u64 j = 0;
|
||||
|
|
@ -225,9 +422,17 @@ static __u64 jent_lfsr_time(struct rand_data *ec, __u64 time, __u64 loop_cnt)
|
|||
new ^= tmp;
|
||||
}
|
||||
}
|
||||
ec->data = new;
|
||||
|
||||
return fold_loop_cnt;
|
||||
/*
|
||||
* If the time stamp is stuck, do not finally insert the value into
|
||||
* the entropy pool. Although this operation should not do any harm
|
||||
* even when the time stamp has no entropy, SP800-90B requires that
|
||||
* any conditioning operation (SP800-90B considers the LFSR to be a
|
||||
* conditioning operation) to have an identical amount of input
|
||||
* data according to section 3.1.5.
|
||||
*/
|
||||
if (!stuck)
|
||||
ec->data = new;
|
||||
}
|
||||
|
||||
/**
|
||||
|
|
@ -248,16 +453,13 @@ static __u64 jent_lfsr_time(struct rand_data *ec, __u64 time, __u64 loop_cnt)
|
|||
* to reliably access either L3 or memory, the ec->mem memory must be quite
|
||||
* large which is usually not desirable.
|
||||
*
|
||||
* Input:
|
||||
* @ec Reference to the entropy collector with the memory access data -- if
|
||||
* the reference to the memory block to be accessed is NULL, this noise
|
||||
* source is disabled
|
||||
* @loop_cnt if a value not equal to 0 is set, use the given value as number of
|
||||
* loops to perform the folding
|
||||
*
|
||||
* @return Number of memory access operations
|
||||
* @ec [in] Reference to the entropy collector with the memory access data -- if
|
||||
* the reference to the memory block to be accessed is NULL, this noise
|
||||
* source is disabled
|
||||
* @loop_cnt [in] if a value not equal to 0 is set, use the given value
|
||||
* number of loops to perform the LFSR
|
||||
*/
|
||||
static unsigned int jent_memaccess(struct rand_data *ec, __u64 loop_cnt)
|
||||
static void jent_memaccess(struct rand_data *ec, __u64 loop_cnt)
|
||||
{
|
||||
unsigned int wrap = 0;
|
||||
__u64 i = 0;
|
||||
|
|
@ -267,7 +469,7 @@ static unsigned int jent_memaccess(struct rand_data *ec, __u64 loop_cnt)
|
|||
jent_loop_shuffle(ec, MAX_ACC_LOOP_BIT, MIN_ACC_LOOP_BIT);
|
||||
|
||||
if (NULL == ec || NULL == ec->mem)
|
||||
return 0;
|
||||
return;
|
||||
wrap = ec->memblocksize * ec->memblocks;
|
||||
|
||||
/*
|
||||
|
|
@ -293,43 +495,11 @@ static unsigned int jent_memaccess(struct rand_data *ec, __u64 loop_cnt)
|
|||
ec->memlocation = ec->memlocation + ec->memblocksize - 1;
|
||||
ec->memlocation = ec->memlocation % wrap;
|
||||
}
|
||||
return i;
|
||||
}
|
||||
|
||||
/***************************************************************************
|
||||
* Start of entropy processing logic
|
||||
***************************************************************************/
|
||||
|
||||
/**
|
||||
* Stuck test by checking the:
|
||||
* 1st derivation of the jitter measurement (time delta)
|
||||
* 2nd derivation of the jitter measurement (delta of time deltas)
|
||||
* 3rd derivation of the jitter measurement (delta of delta of time deltas)
|
||||
*
|
||||
* All values must always be non-zero.
|
||||
*
|
||||
* Input:
|
||||
* @ec Reference to entropy collector
|
||||
* @current_delta Jitter time delta
|
||||
*
|
||||
* @return
|
||||
* 0 jitter measurement not stuck (good bit)
|
||||
* 1 jitter measurement stuck (reject bit)
|
||||
*/
|
||||
static int jent_stuck(struct rand_data *ec, __u64 current_delta)
|
||||
{
|
||||
__s64 delta2 = ec->last_delta - current_delta;
|
||||
__s64 delta3 = delta2 - ec->last_delta2;
|
||||
|
||||
ec->last_delta = current_delta;
|
||||
ec->last_delta2 = delta2;
|
||||
|
||||
if (!current_delta || !delta2 || !delta3)
|
||||
return 1;
|
||||
|
||||
return 0;
|
||||
}
|
||||
|
||||
/**
|
||||
* This is the heart of the entropy generation: calculate time deltas and
|
||||
* use the CPU jitter in the time deltas. The jitter is injected into the
|
||||
|
|
@ -339,8 +509,7 @@ static int jent_stuck(struct rand_data *ec, __u64 current_delta)
|
|||
* of this function! This can be done by calling this function
|
||||
* and not using its result.
|
||||
*
|
||||
* Input:
|
||||
* @entropy_collector Reference to entropy collector
|
||||
* @ec [in] Reference to entropy collector
|
||||
*
|
||||
* @return result of stuck test
|
||||
*/
|
||||
|
|
@ -348,6 +517,7 @@ static int jent_measure_jitter(struct rand_data *ec)
|
|||
{
|
||||
__u64 time = 0;
|
||||
__u64 current_delta = 0;
|
||||
int stuck;
|
||||
|
||||
/* Invoke one noise source before time measurement to add variations */
|
||||
jent_memaccess(ec, 0);
|
||||
|
|
@ -357,22 +527,23 @@ static int jent_measure_jitter(struct rand_data *ec)
|
|||
* invocation to measure the timing variations
|
||||
*/
|
||||
jent_get_nstime(&time);
|
||||
current_delta = time - ec->prev_time;
|
||||
current_delta = jent_delta(ec->prev_time, time);
|
||||
ec->prev_time = time;
|
||||
|
||||
/* Now call the next noise sources which also injects the data */
|
||||
jent_lfsr_time(ec, current_delta, 0);
|
||||
|
||||
/* Check whether we have a stuck measurement. */
|
||||
return jent_stuck(ec, current_delta);
|
||||
stuck = jent_stuck(ec, current_delta);
|
||||
|
||||
/* Now call the next noise sources which also injects the data */
|
||||
jent_lfsr_time(ec, current_delta, 0, stuck);
|
||||
|
||||
return stuck;
|
||||
}
|
||||
|
||||
/**
|
||||
* Generator of one 64 bit random number
|
||||
* Function fills rand_data->data
|
||||
*
|
||||
* Input:
|
||||
* @ec Reference to entropy collector
|
||||
* @ec [in] Reference to entropy collector
|
||||
*/
|
||||
static void jent_gen_entropy(struct rand_data *ec)
|
||||
{
|
||||
|
|
@ -395,31 +566,6 @@ static void jent_gen_entropy(struct rand_data *ec)
|
|||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* The continuous test required by FIPS 140-2 -- the function automatically
|
||||
* primes the test if needed.
|
||||
*
|
||||
* Return:
|
||||
* 0 if FIPS test passed
|
||||
* < 0 if FIPS test failed
|
||||
*/
|
||||
static void jent_fips_test(struct rand_data *ec)
|
||||
{
|
||||
if (!jent_fips_enabled())
|
||||
return;
|
||||
|
||||
/* prime the FIPS test */
|
||||
if (!ec->old_data) {
|
||||
ec->old_data = ec->data;
|
||||
jent_gen_entropy(ec);
|
||||
}
|
||||
|
||||
if (ec->data == ec->old_data)
|
||||
jent_panic("jitterentropy: Duplicate output detected\n");
|
||||
|
||||
ec->old_data = ec->data;
|
||||
}
|
||||
|
||||
/**
|
||||
* Entry function: Obtain entropy for the caller.
|
||||
*
|
||||
|
|
@ -430,17 +576,18 @@ static void jent_fips_test(struct rand_data *ec)
|
|||
* This function truncates the last 64 bit entropy value output to the exact
|
||||
* size specified by the caller.
|
||||
*
|
||||
* Input:
|
||||
* @ec Reference to entropy collector
|
||||
* @data pointer to buffer for storing random data -- buffer must already
|
||||
* exist
|
||||
* @len size of the buffer, specifying also the requested number of random
|
||||
* in bytes
|
||||
* @ec [in] Reference to entropy collector
|
||||
* @data [in] pointer to buffer for storing random data -- buffer must already
|
||||
* exist
|
||||
* @len [in] size of the buffer, specifying also the requested number of random
|
||||
* in bytes
|
||||
*
|
||||
* @return 0 when request is fulfilled or an error
|
||||
*
|
||||
* The following error codes can occur:
|
||||
* -1 entropy_collector is NULL
|
||||
* -2 RCT failed
|
||||
* -3 APT test failed
|
||||
*/
|
||||
int jent_read_entropy(struct rand_data *ec, unsigned char *data,
|
||||
unsigned int len)
|
||||
|
|
@ -454,7 +601,42 @@ int jent_read_entropy(struct rand_data *ec, unsigned char *data,
|
|||
unsigned int tocopy;
|
||||
|
||||
jent_gen_entropy(ec);
|
||||
jent_fips_test(ec);
|
||||
|
||||
if (jent_health_failure(ec)) {
|
||||
int ret;
|
||||
|
||||
if (jent_rct_failure(ec))
|
||||
ret = -2;
|
||||
else
|
||||
ret = -3;
|
||||
|
||||
/*
|
||||
* Re-initialize the noise source
|
||||
*
|
||||
* If the health test fails, the Jitter RNG remains
|
||||
* in failure state and will return a health failure
|
||||
* during next invocation.
|
||||
*/
|
||||
if (jent_entropy_init())
|
||||
return ret;
|
||||
|
||||
/* Set APT to initial state */
|
||||
jent_apt_reset(ec, 0);
|
||||
ec->apt_base_set = 0;
|
||||
|
||||
/* Set RCT to initial state */
|
||||
ec->rct_count = 0;
|
||||
|
||||
/* Re-enable Jitter RNG */
|
||||
ec->health_failure = 0;
|
||||
|
||||
/*
|
||||
* Return the health test failure status to the
|
||||
* caller as the generated value is not appropriate.
|
||||
*/
|
||||
return ret;
|
||||
}
|
||||
|
||||
if ((DATA_SIZE_BITS / 8) < len)
|
||||
tocopy = (DATA_SIZE_BITS / 8);
|
||||
else
|
||||
|
|
@ -518,11 +700,15 @@ int jent_entropy_init(void)
|
|||
int i;
|
||||
__u64 delta_sum = 0;
|
||||
__u64 old_delta = 0;
|
||||
unsigned int nonstuck = 0;
|
||||
int time_backwards = 0;
|
||||
int count_mod = 0;
|
||||
int count_stuck = 0;
|
||||
struct rand_data ec = { 0 };
|
||||
|
||||
/* Required for RCT */
|
||||
ec.osr = 1;
|
||||
|
||||
/* We could perform statistical tests here, but the problem is
|
||||
* that we only have a few loop counts to do testing. These
|
||||
* loop counts may show some slight skew and we produce
|
||||
|
|
@ -544,8 +730,10 @@ int jent_entropy_init(void)
|
|||
/*
|
||||
* TESTLOOPCOUNT needs some loops to identify edge systems. 100 is
|
||||
* definitely too little.
|
||||
*
|
||||
* SP800-90B requires at least 1024 initial test cycles.
|
||||
*/
|
||||
#define TESTLOOPCOUNT 300
|
||||
#define TESTLOOPCOUNT 1024
|
||||
#define CLEARCACHE 100
|
||||
for (i = 0; (TESTLOOPCOUNT + CLEARCACHE) > i; i++) {
|
||||
__u64 time = 0;
|
||||
|
|
@ -557,13 +745,13 @@ int jent_entropy_init(void)
|
|||
/* Invoke core entropy collection logic */
|
||||
jent_get_nstime(&time);
|
||||
ec.prev_time = time;
|
||||
jent_lfsr_time(&ec, time, 0);
|
||||
jent_lfsr_time(&ec, time, 0, 0);
|
||||
jent_get_nstime(&time2);
|
||||
|
||||
/* test whether timer works */
|
||||
if (!time || !time2)
|
||||
return JENT_ENOTIME;
|
||||
delta = time2 - time;
|
||||
delta = jent_delta(time, time2);
|
||||
/*
|
||||
* test whether timer is fine grained enough to provide
|
||||
* delta even when called shortly after each other -- this
|
||||
|
|
@ -586,6 +774,28 @@ int jent_entropy_init(void)
|
|||
|
||||
if (stuck)
|
||||
count_stuck++;
|
||||
else {
|
||||
nonstuck++;
|
||||
|
||||
/*
|
||||
* Ensure that the APT succeeded.
|
||||
*
|
||||
* With the check below that count_stuck must be less
|
||||
* than 10% of the overall generated raw entropy values
|
||||
* it is guaranteed that the APT is invoked at
|
||||
* floor((TESTLOOPCOUNT * 0.9) / 64) == 14 times.
|
||||
*/
|
||||
if ((nonstuck % JENT_APT_WINDOW_SIZE) == 0) {
|
||||
jent_apt_reset(&ec,
|
||||
delta & JENT_APT_WORD_MASK);
|
||||
if (jent_health_failure(&ec))
|
||||
return JENT_EHEALTH;
|
||||
}
|
||||
}
|
||||
|
||||
/* Validate RCT */
|
||||
if (jent_rct_failure(&ec))
|
||||
return JENT_ERCT;
|
||||
|
||||
/* test whether we have an increasing timer */
|
||||
if (!(time2 > time))
|
||||
|
|
|
|||
17
crypto/jitterentropy.h
Normal file
17
crypto/jitterentropy.h
Normal file
|
|
@ -0,0 +1,17 @@
|
|||
// SPDX-License-Identifier: GPL-2.0-or-later
|
||||
|
||||
extern void *jent_zalloc(unsigned int len);
|
||||
extern void jent_zfree(void *ptr);
|
||||
extern int jent_fips_enabled(void);
|
||||
extern void jent_panic(char *s);
|
||||
extern void jent_memcpy(void *dest, const void *src, unsigned int n);
|
||||
extern void jent_get_nstime(__u64 *out);
|
||||
|
||||
struct rand_data;
|
||||
extern int jent_entropy_init(void);
|
||||
extern int jent_read_entropy(struct rand_data *ec, unsigned char *data,
|
||||
unsigned int len);
|
||||
|
||||
extern struct rand_data *jent_entropy_collector_alloc(unsigned int osr,
|
||||
unsigned int flags);
|
||||
extern void jent_entropy_collector_free(struct rand_data *entropy_collector);
|
||||
|
|
@ -33,6 +33,7 @@
|
|||
#include <asm/unaligned.h>
|
||||
#include <crypto/algapi.h>
|
||||
#include <crypto/internal/hash.h>
|
||||
#include <crypto/internal/poly1305.h>
|
||||
#include <crypto/nhpoly1305.h>
|
||||
#include <linux/crypto.h>
|
||||
#include <linux/kernel.h>
|
||||
|
|
@ -78,7 +79,7 @@ static void process_nh_hash_value(struct nhpoly1305_state *state,
|
|||
BUILD_BUG_ON(NH_HASH_BYTES % POLY1305_BLOCK_SIZE != 0);
|
||||
|
||||
poly1305_core_blocks(&state->poly_state, &key->poly_key, state->nh_hash,
|
||||
NH_HASH_BYTES / POLY1305_BLOCK_SIZE);
|
||||
NH_HASH_BYTES / POLY1305_BLOCK_SIZE, 1);
|
||||
}
|
||||
|
||||
/*
|
||||
|
|
@ -209,7 +210,7 @@ int crypto_nhpoly1305_final_helper(struct shash_desc *desc, u8 *dst, nh_t nh_fn)
|
|||
if (state->nh_remaining)
|
||||
process_nh_hash_value(state, key);
|
||||
|
||||
poly1305_core_emit(&state->poly_state, dst);
|
||||
poly1305_core_emit(&state->poly_state, NULL, dst);
|
||||
return 0;
|
||||
}
|
||||
EXPORT_SYMBOL(crypto_nhpoly1305_final_helper);
|
||||
|
|
|
|||
|
|
@ -13,65 +13,33 @@
|
|||
|
||||
#include <crypto/algapi.h>
|
||||
#include <crypto/internal/hash.h>
|
||||
#include <crypto/poly1305.h>
|
||||
#include <crypto/internal/poly1305.h>
|
||||
#include <linux/crypto.h>
|
||||
#include <linux/kernel.h>
|
||||
#include <linux/module.h>
|
||||
#include <asm/unaligned.h>
|
||||
|
||||
static inline u64 mlt(u64 a, u64 b)
|
||||
{
|
||||
return a * b;
|
||||
}
|
||||
|
||||
static inline u32 sr(u64 v, u_char n)
|
||||
{
|
||||
return v >> n;
|
||||
}
|
||||
|
||||
static inline u32 and(u32 v, u32 mask)
|
||||
{
|
||||
return v & mask;
|
||||
}
|
||||
|
||||
int crypto_poly1305_init(struct shash_desc *desc)
|
||||
static int crypto_poly1305_init(struct shash_desc *desc)
|
||||
{
|
||||
struct poly1305_desc_ctx *dctx = shash_desc_ctx(desc);
|
||||
|
||||
poly1305_core_init(&dctx->h);
|
||||
dctx->buflen = 0;
|
||||
dctx->rset = false;
|
||||
dctx->rset = 0;
|
||||
dctx->sset = false;
|
||||
|
||||
return 0;
|
||||
}
|
||||
EXPORT_SYMBOL_GPL(crypto_poly1305_init);
|
||||
|
||||
void poly1305_core_setkey(struct poly1305_key *key, const u8 *raw_key)
|
||||
{
|
||||
/* r &= 0xffffffc0ffffffc0ffffffc0fffffff */
|
||||
key->r[0] = (get_unaligned_le32(raw_key + 0) >> 0) & 0x3ffffff;
|
||||
key->r[1] = (get_unaligned_le32(raw_key + 3) >> 2) & 0x3ffff03;
|
||||
key->r[2] = (get_unaligned_le32(raw_key + 6) >> 4) & 0x3ffc0ff;
|
||||
key->r[3] = (get_unaligned_le32(raw_key + 9) >> 6) & 0x3f03fff;
|
||||
key->r[4] = (get_unaligned_le32(raw_key + 12) >> 8) & 0x00fffff;
|
||||
}
|
||||
EXPORT_SYMBOL_GPL(poly1305_core_setkey);
|
||||
|
||||
/*
|
||||
* Poly1305 requires a unique key for each tag, which implies that we can't set
|
||||
* it on the tfm that gets accessed by multiple users simultaneously. Instead we
|
||||
* expect the key as the first 32 bytes in the update() call.
|
||||
*/
|
||||
unsigned int crypto_poly1305_setdesckey(struct poly1305_desc_ctx *dctx,
|
||||
const u8 *src, unsigned int srclen)
|
||||
static unsigned int crypto_poly1305_setdesckey(struct poly1305_desc_ctx *dctx,
|
||||
const u8 *src, unsigned int srclen)
|
||||
{
|
||||
if (!dctx->sset) {
|
||||
if (!dctx->rset && srclen >= POLY1305_BLOCK_SIZE) {
|
||||
poly1305_core_setkey(&dctx->r, src);
|
||||
poly1305_core_setkey(&dctx->core_r, src);
|
||||
src += POLY1305_BLOCK_SIZE;
|
||||
srclen -= POLY1305_BLOCK_SIZE;
|
||||
dctx->rset = true;
|
||||
dctx->rset = 2;
|
||||
}
|
||||
if (srclen >= POLY1305_BLOCK_SIZE) {
|
||||
dctx->s[0] = get_unaligned_le32(src + 0);
|
||||
|
|
@ -85,86 +53,9 @@ unsigned int crypto_poly1305_setdesckey(struct poly1305_desc_ctx *dctx,
|
|||
}
|
||||
return srclen;
|
||||
}
|
||||
EXPORT_SYMBOL_GPL(crypto_poly1305_setdesckey);
|
||||
|
||||
static void poly1305_blocks_internal(struct poly1305_state *state,
|
||||
const struct poly1305_key *key,
|
||||
const void *src, unsigned int nblocks,
|
||||
u32 hibit)
|
||||
{
|
||||
u32 r0, r1, r2, r3, r4;
|
||||
u32 s1, s2, s3, s4;
|
||||
u32 h0, h1, h2, h3, h4;
|
||||
u64 d0, d1, d2, d3, d4;
|
||||
|
||||
if (!nblocks)
|
||||
return;
|
||||
|
||||
r0 = key->r[0];
|
||||
r1 = key->r[1];
|
||||
r2 = key->r[2];
|
||||
r3 = key->r[3];
|
||||
r4 = key->r[4];
|
||||
|
||||
s1 = r1 * 5;
|
||||
s2 = r2 * 5;
|
||||
s3 = r3 * 5;
|
||||
s4 = r4 * 5;
|
||||
|
||||
h0 = state->h[0];
|
||||
h1 = state->h[1];
|
||||
h2 = state->h[2];
|
||||
h3 = state->h[3];
|
||||
h4 = state->h[4];
|
||||
|
||||
do {
|
||||
/* h += m[i] */
|
||||
h0 += (get_unaligned_le32(src + 0) >> 0) & 0x3ffffff;
|
||||
h1 += (get_unaligned_le32(src + 3) >> 2) & 0x3ffffff;
|
||||
h2 += (get_unaligned_le32(src + 6) >> 4) & 0x3ffffff;
|
||||
h3 += (get_unaligned_le32(src + 9) >> 6) & 0x3ffffff;
|
||||
h4 += (get_unaligned_le32(src + 12) >> 8) | hibit;
|
||||
|
||||
/* h *= r */
|
||||
d0 = mlt(h0, r0) + mlt(h1, s4) + mlt(h2, s3) +
|
||||
mlt(h3, s2) + mlt(h4, s1);
|
||||
d1 = mlt(h0, r1) + mlt(h1, r0) + mlt(h2, s4) +
|
||||
mlt(h3, s3) + mlt(h4, s2);
|
||||
d2 = mlt(h0, r2) + mlt(h1, r1) + mlt(h2, r0) +
|
||||
mlt(h3, s4) + mlt(h4, s3);
|
||||
d3 = mlt(h0, r3) + mlt(h1, r2) + mlt(h2, r1) +
|
||||
mlt(h3, r0) + mlt(h4, s4);
|
||||
d4 = mlt(h0, r4) + mlt(h1, r3) + mlt(h2, r2) +
|
||||
mlt(h3, r1) + mlt(h4, r0);
|
||||
|
||||
/* (partial) h %= p */
|
||||
d1 += sr(d0, 26); h0 = and(d0, 0x3ffffff);
|
||||
d2 += sr(d1, 26); h1 = and(d1, 0x3ffffff);
|
||||
d3 += sr(d2, 26); h2 = and(d2, 0x3ffffff);
|
||||
d4 += sr(d3, 26); h3 = and(d3, 0x3ffffff);
|
||||
h0 += sr(d4, 26) * 5; h4 = and(d4, 0x3ffffff);
|
||||
h1 += h0 >> 26; h0 = h0 & 0x3ffffff;
|
||||
|
||||
src += POLY1305_BLOCK_SIZE;
|
||||
} while (--nblocks);
|
||||
|
||||
state->h[0] = h0;
|
||||
state->h[1] = h1;
|
||||
state->h[2] = h2;
|
||||
state->h[3] = h3;
|
||||
state->h[4] = h4;
|
||||
}
|
||||
|
||||
void poly1305_core_blocks(struct poly1305_state *state,
|
||||
const struct poly1305_key *key,
|
||||
const void *src, unsigned int nblocks)
|
||||
{
|
||||
poly1305_blocks_internal(state, key, src, nblocks, 1 << 24);
|
||||
}
|
||||
EXPORT_SYMBOL_GPL(poly1305_core_blocks);
|
||||
|
||||
static void poly1305_blocks(struct poly1305_desc_ctx *dctx,
|
||||
const u8 *src, unsigned int srclen, u32 hibit)
|
||||
static void poly1305_blocks(struct poly1305_desc_ctx *dctx, const u8 *src,
|
||||
unsigned int srclen)
|
||||
{
|
||||
unsigned int datalen;
|
||||
|
||||
|
|
@ -174,12 +65,12 @@ static void poly1305_blocks(struct poly1305_desc_ctx *dctx,
|
|||
srclen = datalen;
|
||||
}
|
||||
|
||||
poly1305_blocks_internal(&dctx->h, &dctx->r,
|
||||
src, srclen / POLY1305_BLOCK_SIZE, hibit);
|
||||
poly1305_core_blocks(&dctx->h, &dctx->core_r, src,
|
||||
srclen / POLY1305_BLOCK_SIZE, 1);
|
||||
}
|
||||
|
||||
int crypto_poly1305_update(struct shash_desc *desc,
|
||||
const u8 *src, unsigned int srclen)
|
||||
static int crypto_poly1305_update(struct shash_desc *desc,
|
||||
const u8 *src, unsigned int srclen)
|
||||
{
|
||||
struct poly1305_desc_ctx *dctx = shash_desc_ctx(desc);
|
||||
unsigned int bytes;
|
||||
|
|
@ -193,13 +84,13 @@ int crypto_poly1305_update(struct shash_desc *desc,
|
|||
|
||||
if (dctx->buflen == POLY1305_BLOCK_SIZE) {
|
||||
poly1305_blocks(dctx, dctx->buf,
|
||||
POLY1305_BLOCK_SIZE, 1 << 24);
|
||||
POLY1305_BLOCK_SIZE);
|
||||
dctx->buflen = 0;
|
||||
}
|
||||
}
|
||||
|
||||
if (likely(srclen >= POLY1305_BLOCK_SIZE)) {
|
||||
poly1305_blocks(dctx, src, srclen, 1 << 24);
|
||||
poly1305_blocks(dctx, src, srclen);
|
||||
src += srclen - (srclen % POLY1305_BLOCK_SIZE);
|
||||
srclen %= POLY1305_BLOCK_SIZE;
|
||||
}
|
||||
|
|
@ -211,87 +102,17 @@ int crypto_poly1305_update(struct shash_desc *desc,
|
|||
|
||||
return 0;
|
||||
}
|
||||
EXPORT_SYMBOL_GPL(crypto_poly1305_update);
|
||||
|
||||
void poly1305_core_emit(const struct poly1305_state *state, void *dst)
|
||||
{
|
||||
u32 h0, h1, h2, h3, h4;
|
||||
u32 g0, g1, g2, g3, g4;
|
||||
u32 mask;
|
||||
|
||||
/* fully carry h */
|
||||
h0 = state->h[0];
|
||||
h1 = state->h[1];
|
||||
h2 = state->h[2];
|
||||
h3 = state->h[3];
|
||||
h4 = state->h[4];
|
||||
|
||||
h2 += (h1 >> 26); h1 = h1 & 0x3ffffff;
|
||||
h3 += (h2 >> 26); h2 = h2 & 0x3ffffff;
|
||||
h4 += (h3 >> 26); h3 = h3 & 0x3ffffff;
|
||||
h0 += (h4 >> 26) * 5; h4 = h4 & 0x3ffffff;
|
||||
h1 += (h0 >> 26); h0 = h0 & 0x3ffffff;
|
||||
|
||||
/* compute h + -p */
|
||||
g0 = h0 + 5;
|
||||
g1 = h1 + (g0 >> 26); g0 &= 0x3ffffff;
|
||||
g2 = h2 + (g1 >> 26); g1 &= 0x3ffffff;
|
||||
g3 = h3 + (g2 >> 26); g2 &= 0x3ffffff;
|
||||
g4 = h4 + (g3 >> 26) - (1 << 26); g3 &= 0x3ffffff;
|
||||
|
||||
/* select h if h < p, or h + -p if h >= p */
|
||||
mask = (g4 >> ((sizeof(u32) * 8) - 1)) - 1;
|
||||
g0 &= mask;
|
||||
g1 &= mask;
|
||||
g2 &= mask;
|
||||
g3 &= mask;
|
||||
g4 &= mask;
|
||||
mask = ~mask;
|
||||
h0 = (h0 & mask) | g0;
|
||||
h1 = (h1 & mask) | g1;
|
||||
h2 = (h2 & mask) | g2;
|
||||
h3 = (h3 & mask) | g3;
|
||||
h4 = (h4 & mask) | g4;
|
||||
|
||||
/* h = h % (2^128) */
|
||||
put_unaligned_le32((h0 >> 0) | (h1 << 26), dst + 0);
|
||||
put_unaligned_le32((h1 >> 6) | (h2 << 20), dst + 4);
|
||||
put_unaligned_le32((h2 >> 12) | (h3 << 14), dst + 8);
|
||||
put_unaligned_le32((h3 >> 18) | (h4 << 8), dst + 12);
|
||||
}
|
||||
EXPORT_SYMBOL_GPL(poly1305_core_emit);
|
||||
|
||||
int crypto_poly1305_final(struct shash_desc *desc, u8 *dst)
|
||||
static int crypto_poly1305_final(struct shash_desc *desc, u8 *dst)
|
||||
{
|
||||
struct poly1305_desc_ctx *dctx = shash_desc_ctx(desc);
|
||||
__le32 digest[4];
|
||||
u64 f = 0;
|
||||
|
||||
if (unlikely(!dctx->sset))
|
||||
return -ENOKEY;
|
||||
|
||||
if (unlikely(dctx->buflen)) {
|
||||
dctx->buf[dctx->buflen++] = 1;
|
||||
memset(dctx->buf + dctx->buflen, 0,
|
||||
POLY1305_BLOCK_SIZE - dctx->buflen);
|
||||
poly1305_blocks(dctx, dctx->buf, POLY1305_BLOCK_SIZE, 0);
|
||||
}
|
||||
|
||||
poly1305_core_emit(&dctx->h, digest);
|
||||
|
||||
/* mac = (h + s) % (2^128) */
|
||||
f = (f >> 32) + le32_to_cpu(digest[0]) + dctx->s[0];
|
||||
put_unaligned_le32(f, dst + 0);
|
||||
f = (f >> 32) + le32_to_cpu(digest[1]) + dctx->s[1];
|
||||
put_unaligned_le32(f, dst + 4);
|
||||
f = (f >> 32) + le32_to_cpu(digest[2]) + dctx->s[2];
|
||||
put_unaligned_le32(f, dst + 8);
|
||||
f = (f >> 32) + le32_to_cpu(digest[3]) + dctx->s[3];
|
||||
put_unaligned_le32(f, dst + 12);
|
||||
|
||||
poly1305_final_generic(dctx, dst);
|
||||
return 0;
|
||||
}
|
||||
EXPORT_SYMBOL_GPL(crypto_poly1305_final);
|
||||
|
||||
static struct shash_alg poly1305_alg = {
|
||||
.digestsize = POLY1305_DIGEST_SIZE,
|
||||
|
|
|
|||
|
|
@ -35,27 +35,31 @@ EXPORT_SYMBOL_GPL(sha256_zero_message_hash);
|
|||
|
||||
static int crypto_sha256_init(struct shash_desc *desc)
|
||||
{
|
||||
return sha256_init(shash_desc_ctx(desc));
|
||||
sha256_init(shash_desc_ctx(desc));
|
||||
return 0;
|
||||
}
|
||||
|
||||
static int crypto_sha224_init(struct shash_desc *desc)
|
||||
{
|
||||
return sha224_init(shash_desc_ctx(desc));
|
||||
sha224_init(shash_desc_ctx(desc));
|
||||
return 0;
|
||||
}
|
||||
|
||||
int crypto_sha256_update(struct shash_desc *desc, const u8 *data,
|
||||
unsigned int len)
|
||||
{
|
||||
return sha256_update(shash_desc_ctx(desc), data, len);
|
||||
sha256_update(shash_desc_ctx(desc), data, len);
|
||||
return 0;
|
||||
}
|
||||
EXPORT_SYMBOL(crypto_sha256_update);
|
||||
|
||||
static int crypto_sha256_final(struct shash_desc *desc, u8 *out)
|
||||
{
|
||||
if (crypto_shash_digestsize(desc->tfm) == SHA224_DIGEST_SIZE)
|
||||
return sha224_final(shash_desc_ctx(desc), out);
|
||||
sha224_final(shash_desc_ctx(desc), out);
|
||||
else
|
||||
return sha256_final(shash_desc_ctx(desc), out);
|
||||
sha256_final(shash_desc_ctx(desc), out);
|
||||
return 0;
|
||||
}
|
||||
|
||||
int crypto_sha256_finup(struct shash_desc *desc, const u8 *data,
|
||||
|
|
|
|||
|
|
@ -4323,6 +4323,12 @@ static const struct alg_test_desc alg_test_descs[] = {
|
|||
.alg = "cts(cbc(paes))",
|
||||
.test = alg_test_null,
|
||||
.fips_allowed = 1,
|
||||
}, {
|
||||
.alg = "curve25519",
|
||||
.test = alg_test_kpp,
|
||||
.suite = {
|
||||
.kpp = __VECS(curve25519_tv_template)
|
||||
}
|
||||
}, {
|
||||
.alg = "deflate",
|
||||
.test = alg_test_comp,
|
||||
|
|
|
|||
1225
crypto/testmgr.h
1225
crypto/testmgr.h
File diff suppressed because it is too large
Load diff
|
|
@ -71,6 +71,50 @@ config DUMMY
|
|||
To compile this driver as a module, choose M here: the module
|
||||
will be called dummy.
|
||||
|
||||
config WIREGUARD
|
||||
tristate "WireGuard secure network tunnel"
|
||||
depends on NET && INET
|
||||
depends on IPV6 || !IPV6
|
||||
select NET_UDP_TUNNEL
|
||||
select DST_CACHE
|
||||
select CRYPTO
|
||||
select CRYPTO_LIB_CURVE25519
|
||||
select CRYPTO_LIB_CHACHA20POLY1305
|
||||
select CRYPTO_LIB_BLAKE2S
|
||||
select CRYPTO_CHACHA20_X86_64 if X86 && 64BIT
|
||||
select CRYPTO_POLY1305_X86_64 if X86 && 64BIT
|
||||
select CRYPTO_BLAKE2S_X86 if X86 && 64BIT
|
||||
select CRYPTO_CURVE25519_X86 if X86 && 64BIT
|
||||
select ARM_CRYPTO if ARM
|
||||
select ARM64_CRYPTO if ARM64
|
||||
select CRYPTO_CHACHA20_NEON if ARM || (ARM64 && KERNEL_MODE_NEON)
|
||||
select CRYPTO_POLY1305_NEON if ARM64 && KERNEL_MODE_NEON
|
||||
select CRYPTO_POLY1305_ARM if ARM
|
||||
select CRYPTO_BLAKE2S_ARM if ARM
|
||||
select CRYPTO_CURVE25519_NEON if ARM && KERNEL_MODE_NEON
|
||||
select CRYPTO_CHACHA_MIPS if CPU_MIPS32_R2
|
||||
select CRYPTO_POLY1305_MIPS if MIPS
|
||||
help
|
||||
WireGuard is a secure, fast, and easy to use replacement for IPSec
|
||||
that uses modern cryptography and clever networking tricks. It's
|
||||
designed to be fairly general purpose and abstract enough to fit most
|
||||
use cases, while at the same time remaining extremely simple to
|
||||
configure. See www.wireguard.com for more info.
|
||||
|
||||
It's safe to say Y or M here, as the driver is very lightweight and
|
||||
is only in use when an administrator chooses to add an interface.
|
||||
|
||||
config WIREGUARD_DEBUG
|
||||
bool "Debugging checks and verbose messages"
|
||||
depends on WIREGUARD
|
||||
help
|
||||
This will write log messages for handshake and other events
|
||||
that occur for a WireGuard interface. It will also perform some
|
||||
extra validation checks and unit tests at various points. This is
|
||||
only useful for debugging.
|
||||
|
||||
Say N here unless you know what you're doing.
|
||||
|
||||
config EQUALIZER
|
||||
tristate "EQL (serial line load balancing) support"
|
||||
---help---
|
||||
|
|
|
|||
|
|
@ -10,6 +10,7 @@ obj-$(CONFIG_BONDING) += bonding/
|
|||
obj-$(CONFIG_IPVLAN) += ipvlan/
|
||||
obj-$(CONFIG_IPVTAP) += ipvlan/
|
||||
obj-$(CONFIG_DUMMY) += dummy.o
|
||||
obj-$(CONFIG_WIREGUARD) += wireguard/
|
||||
obj-$(CONFIG_EQUALIZER) += eql.o
|
||||
obj-$(CONFIG_IFB) += ifb.o
|
||||
obj-$(CONFIG_MACSEC) += macsec.o
|
||||
|
|
|
|||
17
drivers/net/wireguard/Makefile
Normal file
17
drivers/net/wireguard/Makefile
Normal file
|
|
@ -0,0 +1,17 @@
|
|||
ccflags-y := -D'pr_fmt(fmt)=KBUILD_MODNAME ": " fmt'
|
||||
ccflags-$(CONFIG_WIREGUARD_DEBUG) += -DDEBUG
|
||||
wireguard-y := main.o
|
||||
wireguard-y += noise.o
|
||||
wireguard-y += device.o
|
||||
wireguard-y += peer.o
|
||||
wireguard-y += timers.o
|
||||
wireguard-y += queueing.o
|
||||
wireguard-y += send.o
|
||||
wireguard-y += receive.o
|
||||
wireguard-y += socket.o
|
||||
wireguard-y += peerlookup.o
|
||||
wireguard-y += allowedips.o
|
||||
wireguard-y += ratelimiter.o
|
||||
wireguard-y += cookie.o
|
||||
wireguard-y += netlink.o
|
||||
obj-$(CONFIG_WIREGUARD) := wireguard.o
|
||||
389
drivers/net/wireguard/allowedips.c
Normal file
389
drivers/net/wireguard/allowedips.c
Normal file
|
|
@ -0,0 +1,389 @@
|
|||
// SPDX-License-Identifier: GPL-2.0
|
||||
/*
|
||||
* Copyright (C) 2015-2019 Jason A. Donenfeld <Jason@zx2c4.com>. All Rights Reserved.
|
||||
*/
|
||||
|
||||
#include "allowedips.h"
|
||||
#include "peer.h"
|
||||
|
||||
enum { MAX_ALLOWEDIPS_DEPTH = 129 };
|
||||
|
||||
static struct kmem_cache *node_cache;
|
||||
|
||||
static void swap_endian(u8 *dst, const u8 *src, u8 bits)
|
||||
{
|
||||
if (bits == 32) {
|
||||
*(u32 *)dst = be32_to_cpu(*(const __be32 *)src);
|
||||
} else if (bits == 128) {
|
||||
((u64 *)dst)[0] = get_unaligned_be64(src);
|
||||
((u64 *)dst)[1] = get_unaligned_be64(src + 8);
|
||||
}
|
||||
}
|
||||
|
||||
static void copy_and_assign_cidr(struct allowedips_node *node, const u8 *src,
|
||||
u8 cidr, u8 bits)
|
||||
{
|
||||
node->cidr = cidr;
|
||||
node->bit_at_a = cidr / 8U;
|
||||
#ifdef __LITTLE_ENDIAN
|
||||
node->bit_at_a ^= (bits / 8U - 1U) % 8U;
|
||||
#endif
|
||||
node->bit_at_b = 7U - (cidr % 8U);
|
||||
node->bitlen = bits;
|
||||
memcpy(node->bits, src, bits / 8U);
|
||||
}
|
||||
|
||||
static inline u8 choose(struct allowedips_node *node, const u8 *key)
|
||||
{
|
||||
return (key[node->bit_at_a] >> node->bit_at_b) & 1;
|
||||
}
|
||||
|
||||
static void push_rcu(struct allowedips_node **stack,
|
||||
struct allowedips_node __rcu *p, unsigned int *len)
|
||||
{
|
||||
if (rcu_access_pointer(p)) {
|
||||
if (WARN_ON(IS_ENABLED(DEBUG) && *len >= MAX_ALLOWEDIPS_DEPTH))
|
||||
return;
|
||||
stack[(*len)++] = rcu_dereference_raw(p);
|
||||
}
|
||||
}
|
||||
|
||||
static void node_free_rcu(struct rcu_head *rcu)
|
||||
{
|
||||
kmem_cache_free(node_cache, container_of(rcu, struct allowedips_node, rcu));
|
||||
}
|
||||
|
||||
static void root_free_rcu(struct rcu_head *rcu)
|
||||
{
|
||||
struct allowedips_node *node, *stack[MAX_ALLOWEDIPS_DEPTH] = {
|
||||
container_of(rcu, struct allowedips_node, rcu) };
|
||||
unsigned int len = 1;
|
||||
|
||||
while (len > 0 && (node = stack[--len])) {
|
||||
push_rcu(stack, node->bit[0], &len);
|
||||
push_rcu(stack, node->bit[1], &len);
|
||||
kmem_cache_free(node_cache, node);
|
||||
}
|
||||
}
|
||||
|
||||
static void root_remove_peer_lists(struct allowedips_node *root)
|
||||
{
|
||||
struct allowedips_node *node, *stack[MAX_ALLOWEDIPS_DEPTH] = { root };
|
||||
unsigned int len = 1;
|
||||
|
||||
while (len > 0 && (node = stack[--len])) {
|
||||
push_rcu(stack, node->bit[0], &len);
|
||||
push_rcu(stack, node->bit[1], &len);
|
||||
if (rcu_access_pointer(node->peer))
|
||||
list_del(&node->peer_list);
|
||||
}
|
||||
}
|
||||
|
||||
static unsigned int fls128(u64 a, u64 b)
|
||||
{
|
||||
return a ? fls64(a) + 64U : fls64(b);
|
||||
}
|
||||
|
||||
static u8 common_bits(const struct allowedips_node *node, const u8 *key,
|
||||
u8 bits)
|
||||
{
|
||||
if (bits == 32)
|
||||
return 32U - fls(*(const u32 *)node->bits ^ *(const u32 *)key);
|
||||
else if (bits == 128)
|
||||
return 128U - fls128(
|
||||
*(const u64 *)&node->bits[0] ^ *(const u64 *)&key[0],
|
||||
*(const u64 *)&node->bits[8] ^ *(const u64 *)&key[8]);
|
||||
return 0;
|
||||
}
|
||||
|
||||
static bool prefix_matches(const struct allowedips_node *node, const u8 *key,
|
||||
u8 bits)
|
||||
{
|
||||
/* This could be much faster if it actually just compared the common
|
||||
* bits properly, by precomputing a mask bswap(~0 << (32 - cidr)), and
|
||||
* the rest, but it turns out that common_bits is already super fast on
|
||||
* modern processors, even taking into account the unfortunate bswap.
|
||||
* So, we just inline it like this instead.
|
||||
*/
|
||||
return common_bits(node, key, bits) >= node->cidr;
|
||||
}
|
||||
|
||||
static struct allowedips_node *find_node(struct allowedips_node *trie, u8 bits,
|
||||
const u8 *key)
|
||||
{
|
||||
struct allowedips_node *node = trie, *found = NULL;
|
||||
|
||||
while (node && prefix_matches(node, key, bits)) {
|
||||
if (rcu_access_pointer(node->peer))
|
||||
found = node;
|
||||
if (node->cidr == bits)
|
||||
break;
|
||||
node = rcu_dereference_bh(node->bit[choose(node, key)]);
|
||||
}
|
||||
return found;
|
||||
}
|
||||
|
||||
/* Returns a strong reference to a peer */
|
||||
static struct wg_peer *lookup(struct allowedips_node __rcu *root, u8 bits,
|
||||
const void *be_ip)
|
||||
{
|
||||
/* Aligned so it can be passed to fls/fls64 */
|
||||
u8 ip[16] __aligned(__alignof(u64));
|
||||
struct allowedips_node *node;
|
||||
struct wg_peer *peer = NULL;
|
||||
|
||||
swap_endian(ip, be_ip, bits);
|
||||
|
||||
rcu_read_lock_bh();
|
||||
retry:
|
||||
node = find_node(rcu_dereference_bh(root), bits, ip);
|
||||
if (node) {
|
||||
peer = wg_peer_get_maybe_zero(rcu_dereference_bh(node->peer));
|
||||
if (!peer)
|
||||
goto retry;
|
||||
}
|
||||
rcu_read_unlock_bh();
|
||||
return peer;
|
||||
}
|
||||
|
||||
static bool node_placement(struct allowedips_node __rcu *trie, const u8 *key,
|
||||
u8 cidr, u8 bits, struct allowedips_node **rnode,
|
||||
struct mutex *lock)
|
||||
{
|
||||
struct allowedips_node *node = rcu_dereference_protected(trie, lockdep_is_held(lock));
|
||||
struct allowedips_node *parent = NULL;
|
||||
bool exact = false;
|
||||
|
||||
while (node && node->cidr <= cidr && prefix_matches(node, key, bits)) {
|
||||
parent = node;
|
||||
if (parent->cidr == cidr) {
|
||||
exact = true;
|
||||
break;
|
||||
}
|
||||
node = rcu_dereference_protected(parent->bit[choose(parent, key)], lockdep_is_held(lock));
|
||||
}
|
||||
*rnode = parent;
|
||||
return exact;
|
||||
}
|
||||
|
||||
static inline void connect_node(struct allowedips_node __rcu **parent, u8 bit, struct allowedips_node *node)
|
||||
{
|
||||
node->parent_bit_packed = (unsigned long)parent | bit;
|
||||
rcu_assign_pointer(*parent, node);
|
||||
}
|
||||
|
||||
static inline void choose_and_connect_node(struct allowedips_node *parent, struct allowedips_node *node)
|
||||
{
|
||||
u8 bit = choose(parent, node->bits);
|
||||
connect_node(&parent->bit[bit], bit, node);
|
||||
}
|
||||
|
||||
static int add(struct allowedips_node __rcu **trie, u8 bits, const u8 *key,
|
||||
u8 cidr, struct wg_peer *peer, struct mutex *lock)
|
||||
{
|
||||
struct allowedips_node *node, *parent, *down, *newnode;
|
||||
|
||||
if (unlikely(cidr > bits || !peer))
|
||||
return -EINVAL;
|
||||
|
||||
if (!rcu_access_pointer(*trie)) {
|
||||
node = kmem_cache_zalloc(node_cache, GFP_KERNEL);
|
||||
if (unlikely(!node))
|
||||
return -ENOMEM;
|
||||
RCU_INIT_POINTER(node->peer, peer);
|
||||
list_add_tail(&node->peer_list, &peer->allowedips_list);
|
||||
copy_and_assign_cidr(node, key, cidr, bits);
|
||||
connect_node(trie, 2, node);
|
||||
return 0;
|
||||
}
|
||||
if (node_placement(*trie, key, cidr, bits, &node, lock)) {
|
||||
rcu_assign_pointer(node->peer, peer);
|
||||
list_move_tail(&node->peer_list, &peer->allowedips_list);
|
||||
return 0;
|
||||
}
|
||||
|
||||
newnode = kmem_cache_zalloc(node_cache, GFP_KERNEL);
|
||||
if (unlikely(!newnode))
|
||||
return -ENOMEM;
|
||||
RCU_INIT_POINTER(newnode->peer, peer);
|
||||
list_add_tail(&newnode->peer_list, &peer->allowedips_list);
|
||||
copy_and_assign_cidr(newnode, key, cidr, bits);
|
||||
|
||||
if (!node) {
|
||||
down = rcu_dereference_protected(*trie, lockdep_is_held(lock));
|
||||
} else {
|
||||
const u8 bit = choose(node, key);
|
||||
down = rcu_dereference_protected(node->bit[bit], lockdep_is_held(lock));
|
||||
if (!down) {
|
||||
connect_node(&node->bit[bit], bit, newnode);
|
||||
return 0;
|
||||
}
|
||||
}
|
||||
cidr = min(cidr, common_bits(down, key, bits));
|
||||
parent = node;
|
||||
|
||||
if (newnode->cidr == cidr) {
|
||||
choose_and_connect_node(newnode, down);
|
||||
if (!parent)
|
||||
connect_node(trie, 2, newnode);
|
||||
else
|
||||
choose_and_connect_node(parent, newnode);
|
||||
return 0;
|
||||
}
|
||||
|
||||
node = kmem_cache_zalloc(node_cache, GFP_KERNEL);
|
||||
if (unlikely(!node)) {
|
||||
list_del(&newnode->peer_list);
|
||||
kmem_cache_free(node_cache, newnode);
|
||||
return -ENOMEM;
|
||||
}
|
||||
INIT_LIST_HEAD(&node->peer_list);
|
||||
copy_and_assign_cidr(node, newnode->bits, cidr, bits);
|
||||
|
||||
choose_and_connect_node(node, down);
|
||||
choose_and_connect_node(node, newnode);
|
||||
if (!parent)
|
||||
connect_node(trie, 2, node);
|
||||
else
|
||||
choose_and_connect_node(parent, node);
|
||||
return 0;
|
||||
}
|
||||
|
||||
void wg_allowedips_init(struct allowedips *table)
|
||||
{
|
||||
table->root4 = table->root6 = NULL;
|
||||
table->seq = 1;
|
||||
}
|
||||
|
||||
void wg_allowedips_free(struct allowedips *table, struct mutex *lock)
|
||||
{
|
||||
struct allowedips_node __rcu *old4 = table->root4, *old6 = table->root6;
|
||||
|
||||
++table->seq;
|
||||
RCU_INIT_POINTER(table->root4, NULL);
|
||||
RCU_INIT_POINTER(table->root6, NULL);
|
||||
if (rcu_access_pointer(old4)) {
|
||||
struct allowedips_node *node = rcu_dereference_protected(old4,
|
||||
lockdep_is_held(lock));
|
||||
|
||||
root_remove_peer_lists(node);
|
||||
call_rcu(&node->rcu, root_free_rcu);
|
||||
}
|
||||
if (rcu_access_pointer(old6)) {
|
||||
struct allowedips_node *node = rcu_dereference_protected(old6,
|
||||
lockdep_is_held(lock));
|
||||
|
||||
root_remove_peer_lists(node);
|
||||
call_rcu(&node->rcu, root_free_rcu);
|
||||
}
|
||||
}
|
||||
|
||||
int wg_allowedips_insert_v4(struct allowedips *table, const struct in_addr *ip,
|
||||
u8 cidr, struct wg_peer *peer, struct mutex *lock)
|
||||
{
|
||||
/* Aligned so it can be passed to fls */
|
||||
u8 key[4] __aligned(__alignof(u32));
|
||||
|
||||
++table->seq;
|
||||
swap_endian(key, (const u8 *)ip, 32);
|
||||
return add(&table->root4, 32, key, cidr, peer, lock);
|
||||
}
|
||||
|
||||
int wg_allowedips_insert_v6(struct allowedips *table, const struct in6_addr *ip,
|
||||
u8 cidr, struct wg_peer *peer, struct mutex *lock)
|
||||
{
|
||||
/* Aligned so it can be passed to fls64 */
|
||||
u8 key[16] __aligned(__alignof(u64));
|
||||
|
||||
++table->seq;
|
||||
swap_endian(key, (const u8 *)ip, 128);
|
||||
return add(&table->root6, 128, key, cidr, peer, lock);
|
||||
}
|
||||
|
||||
void wg_allowedips_remove_by_peer(struct allowedips *table,
|
||||
struct wg_peer *peer, struct mutex *lock)
|
||||
{
|
||||
struct allowedips_node *node, *child, **parent_bit, *parent, *tmp;
|
||||
bool free_parent;
|
||||
|
||||
if (list_empty(&peer->allowedips_list))
|
||||
return;
|
||||
++table->seq;
|
||||
list_for_each_entry_safe(node, tmp, &peer->allowedips_list, peer_list) {
|
||||
list_del_init(&node->peer_list);
|
||||
RCU_INIT_POINTER(node->peer, NULL);
|
||||
if (node->bit[0] && node->bit[1])
|
||||
continue;
|
||||
child = rcu_dereference_protected(node->bit[!rcu_access_pointer(node->bit[0])],
|
||||
lockdep_is_held(lock));
|
||||
if (child)
|
||||
child->parent_bit_packed = node->parent_bit_packed;
|
||||
parent_bit = (struct allowedips_node **)(node->parent_bit_packed & ~3UL);
|
||||
*parent_bit = child;
|
||||
parent = (void *)parent_bit -
|
||||
offsetof(struct allowedips_node, bit[node->parent_bit_packed & 1]);
|
||||
free_parent = !rcu_access_pointer(node->bit[0]) &&
|
||||
!rcu_access_pointer(node->bit[1]) &&
|
||||
(node->parent_bit_packed & 3) <= 1 &&
|
||||
!rcu_access_pointer(parent->peer);
|
||||
if (free_parent)
|
||||
child = rcu_dereference_protected(
|
||||
parent->bit[!(node->parent_bit_packed & 1)],
|
||||
lockdep_is_held(lock));
|
||||
call_rcu(&node->rcu, node_free_rcu);
|
||||
if (!free_parent)
|
||||
continue;
|
||||
if (child)
|
||||
child->parent_bit_packed = parent->parent_bit_packed;
|
||||
*(struct allowedips_node **)(parent->parent_bit_packed & ~3UL) = child;
|
||||
call_rcu(&parent->rcu, node_free_rcu);
|
||||
}
|
||||
}
|
||||
|
||||
int wg_allowedips_read_node(struct allowedips_node *node, u8 ip[16], u8 *cidr)
|
||||
{
|
||||
const unsigned int cidr_bytes = DIV_ROUND_UP(node->cidr, 8U);
|
||||
swap_endian(ip, node->bits, node->bitlen);
|
||||
memset(ip + cidr_bytes, 0, node->bitlen / 8U - cidr_bytes);
|
||||
if (node->cidr)
|
||||
ip[cidr_bytes - 1U] &= ~0U << (-node->cidr % 8U);
|
||||
|
||||
*cidr = node->cidr;
|
||||
return node->bitlen == 32 ? AF_INET : AF_INET6;
|
||||
}
|
||||
|
||||
/* Returns a strong reference to a peer */
|
||||
struct wg_peer *wg_allowedips_lookup_dst(struct allowedips *table,
|
||||
struct sk_buff *skb)
|
||||
{
|
||||
if (skb->protocol == htons(ETH_P_IP))
|
||||
return lookup(table->root4, 32, &ip_hdr(skb)->daddr);
|
||||
else if (skb->protocol == htons(ETH_P_IPV6))
|
||||
return lookup(table->root6, 128, &ipv6_hdr(skb)->daddr);
|
||||
return NULL;
|
||||
}
|
||||
|
||||
/* Returns a strong reference to a peer */
|
||||
struct wg_peer *wg_allowedips_lookup_src(struct allowedips *table,
|
||||
struct sk_buff *skb)
|
||||
{
|
||||
if (skb->protocol == htons(ETH_P_IP))
|
||||
return lookup(table->root4, 32, &ip_hdr(skb)->saddr);
|
||||
else if (skb->protocol == htons(ETH_P_IPV6))
|
||||
return lookup(table->root6, 128, &ipv6_hdr(skb)->saddr);
|
||||
return NULL;
|
||||
}
|
||||
|
||||
int __init wg_allowedips_slab_init(void)
|
||||
{
|
||||
node_cache = KMEM_CACHE(allowedips_node, 0);
|
||||
return node_cache ? 0 : -ENOMEM;
|
||||
}
|
||||
|
||||
void wg_allowedips_slab_uninit(void)
|
||||
{
|
||||
rcu_barrier();
|
||||
kmem_cache_destroy(node_cache);
|
||||
}
|
||||
|
||||
#include "selftest/allowedips.c"
|
||||
59
drivers/net/wireguard/allowedips.h
Normal file
59
drivers/net/wireguard/allowedips.h
Normal file
|
|
@ -0,0 +1,59 @@
|
|||
/* SPDX-License-Identifier: GPL-2.0 */
|
||||
/*
|
||||
* Copyright (C) 2015-2019 Jason A. Donenfeld <Jason@zx2c4.com>. All Rights Reserved.
|
||||
*/
|
||||
|
||||
#ifndef _WG_ALLOWEDIPS_H
|
||||
#define _WG_ALLOWEDIPS_H
|
||||
|
||||
#include <linux/mutex.h>
|
||||
#include <linux/ip.h>
|
||||
#include <linux/ipv6.h>
|
||||
|
||||
struct wg_peer;
|
||||
|
||||
struct allowedips_node {
|
||||
struct wg_peer __rcu *peer;
|
||||
struct allowedips_node __rcu *bit[2];
|
||||
u8 cidr, bit_at_a, bit_at_b, bitlen;
|
||||
u8 bits[16] __aligned(__alignof(u64));
|
||||
|
||||
/* Keep rarely used members at bottom to be beyond cache line. */
|
||||
unsigned long parent_bit_packed;
|
||||
union {
|
||||
struct list_head peer_list;
|
||||
struct rcu_head rcu;
|
||||
};
|
||||
};
|
||||
|
||||
struct allowedips {
|
||||
struct allowedips_node __rcu *root4;
|
||||
struct allowedips_node __rcu *root6;
|
||||
u64 seq;
|
||||
} __aligned(4); /* We pack the lower 2 bits of &root, but m68k only gives 16-bit alignment. */
|
||||
|
||||
void wg_allowedips_init(struct allowedips *table);
|
||||
void wg_allowedips_free(struct allowedips *table, struct mutex *mutex);
|
||||
int wg_allowedips_insert_v4(struct allowedips *table, const struct in_addr *ip,
|
||||
u8 cidr, struct wg_peer *peer, struct mutex *lock);
|
||||
int wg_allowedips_insert_v6(struct allowedips *table, const struct in6_addr *ip,
|
||||
u8 cidr, struct wg_peer *peer, struct mutex *lock);
|
||||
void wg_allowedips_remove_by_peer(struct allowedips *table,
|
||||
struct wg_peer *peer, struct mutex *lock);
|
||||
/* The ip input pointer should be __aligned(__alignof(u64))) */
|
||||
int wg_allowedips_read_node(struct allowedips_node *node, u8 ip[16], u8 *cidr);
|
||||
|
||||
/* These return a strong reference to a peer: */
|
||||
struct wg_peer *wg_allowedips_lookup_dst(struct allowedips *table,
|
||||
struct sk_buff *skb);
|
||||
struct wg_peer *wg_allowedips_lookup_src(struct allowedips *table,
|
||||
struct sk_buff *skb);
|
||||
|
||||
#ifdef DEBUG
|
||||
bool wg_allowedips_selftest(void);
|
||||
#endif
|
||||
|
||||
int wg_allowedips_slab_init(void);
|
||||
void wg_allowedips_slab_uninit(void);
|
||||
|
||||
#endif /* _WG_ALLOWEDIPS_H */
|
||||
236
drivers/net/wireguard/cookie.c
Normal file
236
drivers/net/wireguard/cookie.c
Normal file
|
|
@ -0,0 +1,236 @@
|
|||
// SPDX-License-Identifier: GPL-2.0
|
||||
/*
|
||||
* Copyright (C) 2015-2019 Jason A. Donenfeld <Jason@zx2c4.com>. All Rights Reserved.
|
||||
*/
|
||||
|
||||
#include "cookie.h"
|
||||
#include "peer.h"
|
||||
#include "device.h"
|
||||
#include "messages.h"
|
||||
#include "ratelimiter.h"
|
||||
#include "timers.h"
|
||||
|
||||
#include <crypto/blake2s.h>
|
||||
#include <crypto/chacha20poly1305.h>
|
||||
|
||||
#include <net/ipv6.h>
|
||||
#include <crypto/algapi.h>
|
||||
|
||||
void wg_cookie_checker_init(struct cookie_checker *checker,
|
||||
struct wg_device *wg)
|
||||
{
|
||||
init_rwsem(&checker->secret_lock);
|
||||
checker->secret_birthdate = ktime_get_coarse_boottime_ns();
|
||||
get_random_bytes(checker->secret, NOISE_HASH_LEN);
|
||||
checker->device = wg;
|
||||
}
|
||||
|
||||
enum { COOKIE_KEY_LABEL_LEN = 8 };
|
||||
static const u8 mac1_key_label[COOKIE_KEY_LABEL_LEN] = "mac1----";
|
||||
static const u8 cookie_key_label[COOKIE_KEY_LABEL_LEN] = "cookie--";
|
||||
|
||||
static void precompute_key(u8 key[NOISE_SYMMETRIC_KEY_LEN],
|
||||
const u8 pubkey[NOISE_PUBLIC_KEY_LEN],
|
||||
const u8 label[COOKIE_KEY_LABEL_LEN])
|
||||
{
|
||||
struct blake2s_state blake;
|
||||
|
||||
blake2s_init(&blake, NOISE_SYMMETRIC_KEY_LEN);
|
||||
blake2s_update(&blake, label, COOKIE_KEY_LABEL_LEN);
|
||||
blake2s_update(&blake, pubkey, NOISE_PUBLIC_KEY_LEN);
|
||||
blake2s_final(&blake, key);
|
||||
}
|
||||
|
||||
/* Must hold peer->handshake.static_identity->lock */
|
||||
void wg_cookie_checker_precompute_device_keys(struct cookie_checker *checker)
|
||||
{
|
||||
if (likely(checker->device->static_identity.has_identity)) {
|
||||
precompute_key(checker->cookie_encryption_key,
|
||||
checker->device->static_identity.static_public,
|
||||
cookie_key_label);
|
||||
precompute_key(checker->message_mac1_key,
|
||||
checker->device->static_identity.static_public,
|
||||
mac1_key_label);
|
||||
} else {
|
||||
memset(checker->cookie_encryption_key, 0,
|
||||
NOISE_SYMMETRIC_KEY_LEN);
|
||||
memset(checker->message_mac1_key, 0, NOISE_SYMMETRIC_KEY_LEN);
|
||||
}
|
||||
}
|
||||
|
||||
void wg_cookie_checker_precompute_peer_keys(struct wg_peer *peer)
|
||||
{
|
||||
precompute_key(peer->latest_cookie.cookie_decryption_key,
|
||||
peer->handshake.remote_static, cookie_key_label);
|
||||
precompute_key(peer->latest_cookie.message_mac1_key,
|
||||
peer->handshake.remote_static, mac1_key_label);
|
||||
}
|
||||
|
||||
void wg_cookie_init(struct cookie *cookie)
|
||||
{
|
||||
memset(cookie, 0, sizeof(*cookie));
|
||||
init_rwsem(&cookie->lock);
|
||||
}
|
||||
|
||||
static void compute_mac1(u8 mac1[COOKIE_LEN], const void *message, size_t len,
|
||||
const u8 key[NOISE_SYMMETRIC_KEY_LEN])
|
||||
{
|
||||
len = len - sizeof(struct message_macs) +
|
||||
offsetof(struct message_macs, mac1);
|
||||
blake2s(mac1, message, key, COOKIE_LEN, len, NOISE_SYMMETRIC_KEY_LEN);
|
||||
}
|
||||
|
||||
static void compute_mac2(u8 mac2[COOKIE_LEN], const void *message, size_t len,
|
||||
const u8 cookie[COOKIE_LEN])
|
||||
{
|
||||
len = len - sizeof(struct message_macs) +
|
||||
offsetof(struct message_macs, mac2);
|
||||
blake2s(mac2, message, cookie, COOKIE_LEN, len, COOKIE_LEN);
|
||||
}
|
||||
|
||||
static void make_cookie(u8 cookie[COOKIE_LEN], struct sk_buff *skb,
|
||||
struct cookie_checker *checker)
|
||||
{
|
||||
struct blake2s_state state;
|
||||
|
||||
if (wg_birthdate_has_expired(checker->secret_birthdate,
|
||||
COOKIE_SECRET_MAX_AGE)) {
|
||||
down_write(&checker->secret_lock);
|
||||
checker->secret_birthdate = ktime_get_coarse_boottime_ns();
|
||||
get_random_bytes(checker->secret, NOISE_HASH_LEN);
|
||||
up_write(&checker->secret_lock);
|
||||
}
|
||||
|
||||
down_read(&checker->secret_lock);
|
||||
|
||||
blake2s_init_key(&state, COOKIE_LEN, checker->secret, NOISE_HASH_LEN);
|
||||
if (skb->protocol == htons(ETH_P_IP))
|
||||
blake2s_update(&state, (u8 *)&ip_hdr(skb)->saddr,
|
||||
sizeof(struct in_addr));
|
||||
else if (skb->protocol == htons(ETH_P_IPV6))
|
||||
blake2s_update(&state, (u8 *)&ipv6_hdr(skb)->saddr,
|
||||
sizeof(struct in6_addr));
|
||||
blake2s_update(&state, (u8 *)&udp_hdr(skb)->source, sizeof(__be16));
|
||||
blake2s_final(&state, cookie);
|
||||
|
||||
up_read(&checker->secret_lock);
|
||||
}
|
||||
|
||||
enum cookie_mac_state wg_cookie_validate_packet(struct cookie_checker *checker,
|
||||
struct sk_buff *skb,
|
||||
bool check_cookie)
|
||||
{
|
||||
struct message_macs *macs = (struct message_macs *)
|
||||
(skb->data + skb->len - sizeof(*macs));
|
||||
enum cookie_mac_state ret;
|
||||
u8 computed_mac[COOKIE_LEN];
|
||||
u8 cookie[COOKIE_LEN];
|
||||
|
||||
ret = INVALID_MAC;
|
||||
compute_mac1(computed_mac, skb->data, skb->len,
|
||||
checker->message_mac1_key);
|
||||
if (crypto_memneq(computed_mac, macs->mac1, COOKIE_LEN))
|
||||
goto out;
|
||||
|
||||
ret = VALID_MAC_BUT_NO_COOKIE;
|
||||
|
||||
if (!check_cookie)
|
||||
goto out;
|
||||
|
||||
make_cookie(cookie, skb, checker);
|
||||
|
||||
compute_mac2(computed_mac, skb->data, skb->len, cookie);
|
||||
if (crypto_memneq(computed_mac, macs->mac2, COOKIE_LEN))
|
||||
goto out;
|
||||
|
||||
ret = VALID_MAC_WITH_COOKIE_BUT_RATELIMITED;
|
||||
if (!wg_ratelimiter_allow(skb, dev_net(checker->device->dev)))
|
||||
goto out;
|
||||
|
||||
ret = VALID_MAC_WITH_COOKIE;
|
||||
|
||||
out:
|
||||
return ret;
|
||||
}
|
||||
|
||||
void wg_cookie_add_mac_to_packet(void *message, size_t len,
|
||||
struct wg_peer *peer)
|
||||
{
|
||||
struct message_macs *macs = (struct message_macs *)
|
||||
((u8 *)message + len - sizeof(*macs));
|
||||
|
||||
down_write(&peer->latest_cookie.lock);
|
||||
compute_mac1(macs->mac1, message, len,
|
||||
peer->latest_cookie.message_mac1_key);
|
||||
memcpy(peer->latest_cookie.last_mac1_sent, macs->mac1, COOKIE_LEN);
|
||||
peer->latest_cookie.have_sent_mac1 = true;
|
||||
up_write(&peer->latest_cookie.lock);
|
||||
|
||||
down_read(&peer->latest_cookie.lock);
|
||||
if (peer->latest_cookie.is_valid &&
|
||||
!wg_birthdate_has_expired(peer->latest_cookie.birthdate,
|
||||
COOKIE_SECRET_MAX_AGE - COOKIE_SECRET_LATENCY))
|
||||
compute_mac2(macs->mac2, message, len,
|
||||
peer->latest_cookie.cookie);
|
||||
else
|
||||
memset(macs->mac2, 0, COOKIE_LEN);
|
||||
up_read(&peer->latest_cookie.lock);
|
||||
}
|
||||
|
||||
void wg_cookie_message_create(struct message_handshake_cookie *dst,
|
||||
struct sk_buff *skb, __le32 index,
|
||||
struct cookie_checker *checker)
|
||||
{
|
||||
struct message_macs *macs = (struct message_macs *)
|
||||
((u8 *)skb->data + skb->len - sizeof(*macs));
|
||||
u8 cookie[COOKIE_LEN];
|
||||
|
||||
dst->header.type = cpu_to_le32(MESSAGE_HANDSHAKE_COOKIE);
|
||||
dst->receiver_index = index;
|
||||
get_random_bytes_wait(dst->nonce, COOKIE_NONCE_LEN);
|
||||
|
||||
make_cookie(cookie, skb, checker);
|
||||
xchacha20poly1305_encrypt(dst->encrypted_cookie, cookie, COOKIE_LEN,
|
||||
macs->mac1, COOKIE_LEN, dst->nonce,
|
||||
checker->cookie_encryption_key);
|
||||
}
|
||||
|
||||
void wg_cookie_message_consume(struct message_handshake_cookie *src,
|
||||
struct wg_device *wg)
|
||||
{
|
||||
struct wg_peer *peer = NULL;
|
||||
u8 cookie[COOKIE_LEN];
|
||||
bool ret;
|
||||
|
||||
if (unlikely(!wg_index_hashtable_lookup(wg->index_hashtable,
|
||||
INDEX_HASHTABLE_HANDSHAKE |
|
||||
INDEX_HASHTABLE_KEYPAIR,
|
||||
src->receiver_index, &peer)))
|
||||
return;
|
||||
|
||||
down_read(&peer->latest_cookie.lock);
|
||||
if (unlikely(!peer->latest_cookie.have_sent_mac1)) {
|
||||
up_read(&peer->latest_cookie.lock);
|
||||
goto out;
|
||||
}
|
||||
ret = xchacha20poly1305_decrypt(
|
||||
cookie, src->encrypted_cookie, sizeof(src->encrypted_cookie),
|
||||
peer->latest_cookie.last_mac1_sent, COOKIE_LEN, src->nonce,
|
||||
peer->latest_cookie.cookie_decryption_key);
|
||||
up_read(&peer->latest_cookie.lock);
|
||||
|
||||
if (ret) {
|
||||
down_write(&peer->latest_cookie.lock);
|
||||
memcpy(peer->latest_cookie.cookie, cookie, COOKIE_LEN);
|
||||
peer->latest_cookie.birthdate = ktime_get_coarse_boottime_ns();
|
||||
peer->latest_cookie.is_valid = true;
|
||||
peer->latest_cookie.have_sent_mac1 = false;
|
||||
up_write(&peer->latest_cookie.lock);
|
||||
} else {
|
||||
net_dbg_ratelimited("%s: Could not decrypt invalid cookie response\n",
|
||||
wg->dev->name);
|
||||
}
|
||||
|
||||
out:
|
||||
wg_peer_put(peer);
|
||||
}
|
||||
59
drivers/net/wireguard/cookie.h
Normal file
59
drivers/net/wireguard/cookie.h
Normal file
|
|
@ -0,0 +1,59 @@
|
|||
/* SPDX-License-Identifier: GPL-2.0 */
|
||||
/*
|
||||
* Copyright (C) 2015-2019 Jason A. Donenfeld <Jason@zx2c4.com>. All Rights Reserved.
|
||||
*/
|
||||
|
||||
#ifndef _WG_COOKIE_H
|
||||
#define _WG_COOKIE_H
|
||||
|
||||
#include "messages.h"
|
||||
#include <linux/rwsem.h>
|
||||
|
||||
struct wg_peer;
|
||||
|
||||
struct cookie_checker {
|
||||
u8 secret[NOISE_HASH_LEN];
|
||||
u8 cookie_encryption_key[NOISE_SYMMETRIC_KEY_LEN];
|
||||
u8 message_mac1_key[NOISE_SYMMETRIC_KEY_LEN];
|
||||
u64 secret_birthdate;
|
||||
struct rw_semaphore secret_lock;
|
||||
struct wg_device *device;
|
||||
};
|
||||
|
||||
struct cookie {
|
||||
u64 birthdate;
|
||||
bool is_valid;
|
||||
u8 cookie[COOKIE_LEN];
|
||||
bool have_sent_mac1;
|
||||
u8 last_mac1_sent[COOKIE_LEN];
|
||||
u8 cookie_decryption_key[NOISE_SYMMETRIC_KEY_LEN];
|
||||
u8 message_mac1_key[NOISE_SYMMETRIC_KEY_LEN];
|
||||
struct rw_semaphore lock;
|
||||
};
|
||||
|
||||
enum cookie_mac_state {
|
||||
INVALID_MAC,
|
||||
VALID_MAC_BUT_NO_COOKIE,
|
||||
VALID_MAC_WITH_COOKIE_BUT_RATELIMITED,
|
||||
VALID_MAC_WITH_COOKIE
|
||||
};
|
||||
|
||||
void wg_cookie_checker_init(struct cookie_checker *checker,
|
||||
struct wg_device *wg);
|
||||
void wg_cookie_checker_precompute_device_keys(struct cookie_checker *checker);
|
||||
void wg_cookie_checker_precompute_peer_keys(struct wg_peer *peer);
|
||||
void wg_cookie_init(struct cookie *cookie);
|
||||
|
||||
enum cookie_mac_state wg_cookie_validate_packet(struct cookie_checker *checker,
|
||||
struct sk_buff *skb,
|
||||
bool check_cookie);
|
||||
void wg_cookie_add_mac_to_packet(void *message, size_t len,
|
||||
struct wg_peer *peer);
|
||||
|
||||
void wg_cookie_message_create(struct message_handshake_cookie *src,
|
||||
struct sk_buff *skb, __le32 index,
|
||||
struct cookie_checker *checker);
|
||||
void wg_cookie_message_consume(struct message_handshake_cookie *src,
|
||||
struct wg_device *wg);
|
||||
|
||||
#endif /* _WG_COOKIE_H */
|
||||
461
drivers/net/wireguard/device.c
Normal file
461
drivers/net/wireguard/device.c
Normal file
|
|
@ -0,0 +1,461 @@
|
|||
// SPDX-License-Identifier: GPL-2.0
|
||||
/*
|
||||
* Copyright (C) 2015-2019 Jason A. Donenfeld <Jason@zx2c4.com>. All Rights Reserved.
|
||||
*/
|
||||
|
||||
#include "queueing.h"
|
||||
#include "socket.h"
|
||||
#include "timers.h"
|
||||
#include "device.h"
|
||||
#include "ratelimiter.h"
|
||||
#include "peer.h"
|
||||
#include "messages.h"
|
||||
|
||||
#include <linux/module.h>
|
||||
#include <linux/rtnetlink.h>
|
||||
#include <linux/inet.h>
|
||||
#include <linux/netdevice.h>
|
||||
#include <linux/inetdevice.h>
|
||||
#include <linux/if_arp.h>
|
||||
#include <linux/icmp.h>
|
||||
#include <linux/suspend.h>
|
||||
#include <net/dst_metadata.h>
|
||||
#include <net/icmp.h>
|
||||
#include <net/rtnetlink.h>
|
||||
#include <net/ip_tunnels.h>
|
||||
#include <net/addrconf.h>
|
||||
|
||||
static LIST_HEAD(device_list);
|
||||
|
||||
static int wg_open(struct net_device *dev)
|
||||
{
|
||||
struct in_device *dev_v4 = __in_dev_get_rtnl(dev);
|
||||
struct inet6_dev *dev_v6 = __in6_dev_get(dev);
|
||||
struct wg_device *wg = netdev_priv(dev);
|
||||
struct wg_peer *peer;
|
||||
int ret;
|
||||
|
||||
if (dev_v4) {
|
||||
/* At some point we might put this check near the ip_rt_send_
|
||||
* redirect call of ip_forward in net/ipv4/ip_forward.c, similar
|
||||
* to the current secpath check.
|
||||
*/
|
||||
IN_DEV_CONF_SET(dev_v4, SEND_REDIRECTS, false);
|
||||
IPV4_DEVCONF_ALL(dev_net(dev), SEND_REDIRECTS) = false;
|
||||
}
|
||||
if (dev_v6)
|
||||
dev_v6->cnf.addr_gen_mode = IN6_ADDR_GEN_MODE_NONE;
|
||||
|
||||
mutex_lock(&wg->device_update_lock);
|
||||
ret = wg_socket_init(wg, wg->incoming_port);
|
||||
if (ret < 0)
|
||||
goto out;
|
||||
list_for_each_entry(peer, &wg->peer_list, peer_list) {
|
||||
wg_packet_send_staged_packets(peer);
|
||||
if (peer->persistent_keepalive_interval)
|
||||
wg_packet_send_keepalive(peer);
|
||||
}
|
||||
out:
|
||||
mutex_unlock(&wg->device_update_lock);
|
||||
return ret;
|
||||
}
|
||||
|
||||
#ifdef CONFIG_PM_SLEEP
|
||||
static int wg_pm_notification(struct notifier_block *nb, unsigned long action,
|
||||
void *data)
|
||||
{
|
||||
struct wg_device *wg;
|
||||
struct wg_peer *peer;
|
||||
|
||||
/* If the machine is constantly suspending and resuming, as part of
|
||||
* its normal operation rather than as a somewhat rare event, then we
|
||||
* don't actually want to clear keys.
|
||||
*/
|
||||
if (IS_ENABLED(CONFIG_PM_AUTOSLEEP) || IS_ENABLED(CONFIG_ANDROID))
|
||||
return 0;
|
||||
|
||||
if (action != PM_HIBERNATION_PREPARE && action != PM_SUSPEND_PREPARE)
|
||||
return 0;
|
||||
|
||||
rtnl_lock();
|
||||
list_for_each_entry(wg, &device_list, device_list) {
|
||||
mutex_lock(&wg->device_update_lock);
|
||||
list_for_each_entry(peer, &wg->peer_list, peer_list) {
|
||||
del_timer(&peer->timer_zero_key_material);
|
||||
wg_noise_handshake_clear(&peer->handshake);
|
||||
wg_noise_keypairs_clear(&peer->keypairs);
|
||||
}
|
||||
mutex_unlock(&wg->device_update_lock);
|
||||
}
|
||||
rtnl_unlock();
|
||||
rcu_barrier();
|
||||
return 0;
|
||||
}
|
||||
|
||||
static struct notifier_block pm_notifier = { .notifier_call = wg_pm_notification };
|
||||
#endif
|
||||
|
||||
static int wg_stop(struct net_device *dev)
|
||||
{
|
||||
struct wg_device *wg = netdev_priv(dev);
|
||||
struct wg_peer *peer;
|
||||
struct sk_buff *skb;
|
||||
|
||||
mutex_lock(&wg->device_update_lock);
|
||||
list_for_each_entry(peer, &wg->peer_list, peer_list) {
|
||||
wg_packet_purge_staged_packets(peer);
|
||||
wg_timers_stop(peer);
|
||||
wg_noise_handshake_clear(&peer->handshake);
|
||||
wg_noise_keypairs_clear(&peer->keypairs);
|
||||
wg_noise_reset_last_sent_handshake(&peer->last_sent_handshake);
|
||||
}
|
||||
mutex_unlock(&wg->device_update_lock);
|
||||
while ((skb = ptr_ring_consume(&wg->handshake_queue.ring)) != NULL)
|
||||
kfree_skb(skb);
|
||||
atomic_set(&wg->handshake_queue_len, 0);
|
||||
wg_socket_reinit(wg, NULL, NULL);
|
||||
return 0;
|
||||
}
|
||||
|
||||
static netdev_tx_t wg_xmit(struct sk_buff *skb, struct net_device *dev)
|
||||
{
|
||||
struct wg_device *wg = netdev_priv(dev);
|
||||
struct sk_buff_head packets;
|
||||
struct wg_peer *peer;
|
||||
struct sk_buff *next;
|
||||
sa_family_t family;
|
||||
u32 mtu;
|
||||
int ret;
|
||||
|
||||
if (unlikely(!wg_check_packet_protocol(skb))) {
|
||||
ret = -EPROTONOSUPPORT;
|
||||
net_dbg_ratelimited("%s: Invalid IP packet\n", dev->name);
|
||||
goto err;
|
||||
}
|
||||
|
||||
peer = wg_allowedips_lookup_dst(&wg->peer_allowedips, skb);
|
||||
if (unlikely(!peer)) {
|
||||
ret = -ENOKEY;
|
||||
if (skb->protocol == htons(ETH_P_IP))
|
||||
net_dbg_ratelimited("%s: No peer has allowed IPs matching %pI4\n",
|
||||
dev->name, &ip_hdr(skb)->daddr);
|
||||
else if (skb->protocol == htons(ETH_P_IPV6))
|
||||
net_dbg_ratelimited("%s: No peer has allowed IPs matching %pI6\n",
|
||||
dev->name, &ipv6_hdr(skb)->daddr);
|
||||
goto err_icmp;
|
||||
}
|
||||
|
||||
family = READ_ONCE(peer->endpoint.addr.sa_family);
|
||||
if (unlikely(family != AF_INET && family != AF_INET6)) {
|
||||
ret = -EDESTADDRREQ;
|
||||
net_dbg_ratelimited("%s: No valid endpoint has been configured or discovered for peer %llu\n",
|
||||
dev->name, peer->internal_id);
|
||||
goto err_peer;
|
||||
}
|
||||
|
||||
mtu = skb_valid_dst(skb) ? dst_mtu(skb_dst(skb)) : dev->mtu;
|
||||
|
||||
__skb_queue_head_init(&packets);
|
||||
if (!skb_is_gso(skb)) {
|
||||
skb_mark_not_on_list(skb);
|
||||
} else {
|
||||
struct sk_buff *segs = skb_gso_segment(skb, 0);
|
||||
|
||||
if (unlikely(IS_ERR(segs))) {
|
||||
ret = PTR_ERR(segs);
|
||||
goto err_peer;
|
||||
}
|
||||
dev_kfree_skb(skb);
|
||||
skb = segs;
|
||||
}
|
||||
|
||||
skb_list_walk_safe(skb, skb, next) {
|
||||
skb_mark_not_on_list(skb);
|
||||
|
||||
skb = skb_share_check(skb, GFP_ATOMIC);
|
||||
if (unlikely(!skb))
|
||||
continue;
|
||||
|
||||
/* We only need to keep the original dst around for icmp,
|
||||
* so at this point we're in a position to drop it.
|
||||
*/
|
||||
skb_dst_drop(skb);
|
||||
|
||||
PACKET_CB(skb)->mtu = mtu;
|
||||
|
||||
__skb_queue_tail(&packets, skb);
|
||||
}
|
||||
|
||||
spin_lock_bh(&peer->staged_packet_queue.lock);
|
||||
/* If the queue is getting too big, we start removing the oldest packets
|
||||
* until it's small again. We do this before adding the new packet, so
|
||||
* we don't remove GSO segments that are in excess.
|
||||
*/
|
||||
while (skb_queue_len(&peer->staged_packet_queue) > MAX_STAGED_PACKETS) {
|
||||
dev_kfree_skb(__skb_dequeue(&peer->staged_packet_queue));
|
||||
++dev->stats.tx_dropped;
|
||||
}
|
||||
skb_queue_splice_tail(&packets, &peer->staged_packet_queue);
|
||||
spin_unlock_bh(&peer->staged_packet_queue.lock);
|
||||
|
||||
wg_packet_send_staged_packets(peer);
|
||||
|
||||
wg_peer_put(peer);
|
||||
return NETDEV_TX_OK;
|
||||
|
||||
err_peer:
|
||||
wg_peer_put(peer);
|
||||
err_icmp:
|
||||
if (skb->protocol == htons(ETH_P_IP))
|
||||
icmp_ndo_send(skb, ICMP_DEST_UNREACH, ICMP_HOST_UNREACH, 0);
|
||||
else if (skb->protocol == htons(ETH_P_IPV6))
|
||||
icmpv6_ndo_send(skb, ICMPV6_DEST_UNREACH, ICMPV6_ADDR_UNREACH, 0);
|
||||
err:
|
||||
++dev->stats.tx_errors;
|
||||
kfree_skb(skb);
|
||||
return ret;
|
||||
}
|
||||
|
||||
static const struct net_device_ops netdev_ops = {
|
||||
.ndo_open = wg_open,
|
||||
.ndo_stop = wg_stop,
|
||||
.ndo_start_xmit = wg_xmit,
|
||||
.ndo_get_stats64 = ip_tunnel_get_stats64
|
||||
};
|
||||
|
||||
static void wg_destruct(struct net_device *dev)
|
||||
{
|
||||
struct wg_device *wg = netdev_priv(dev);
|
||||
|
||||
rtnl_lock();
|
||||
list_del(&wg->device_list);
|
||||
rtnl_unlock();
|
||||
mutex_lock(&wg->device_update_lock);
|
||||
rcu_assign_pointer(wg->creating_net, NULL);
|
||||
wg->incoming_port = 0;
|
||||
wg_socket_reinit(wg, NULL, NULL);
|
||||
/* The final references are cleared in the below calls to destroy_workqueue. */
|
||||
wg_peer_remove_all(wg);
|
||||
destroy_workqueue(wg->handshake_receive_wq);
|
||||
destroy_workqueue(wg->handshake_send_wq);
|
||||
destroy_workqueue(wg->packet_crypt_wq);
|
||||
wg_packet_queue_free(&wg->handshake_queue, true);
|
||||
wg_packet_queue_free(&wg->decrypt_queue, false);
|
||||
wg_packet_queue_free(&wg->encrypt_queue, false);
|
||||
rcu_barrier(); /* Wait for all the peers to be actually freed. */
|
||||
wg_ratelimiter_uninit();
|
||||
memzero_explicit(&wg->static_identity, sizeof(wg->static_identity));
|
||||
free_percpu(dev->tstats);
|
||||
kvfree(wg->index_hashtable);
|
||||
kvfree(wg->peer_hashtable);
|
||||
mutex_unlock(&wg->device_update_lock);
|
||||
|
||||
pr_debug("%s: Interface destroyed\n", dev->name);
|
||||
free_netdev(dev);
|
||||
}
|
||||
|
||||
static const struct device_type device_type = { .name = KBUILD_MODNAME };
|
||||
|
||||
static void wg_setup(struct net_device *dev)
|
||||
{
|
||||
struct wg_device *wg = netdev_priv(dev);
|
||||
enum { WG_NETDEV_FEATURES = NETIF_F_HW_CSUM | NETIF_F_RXCSUM |
|
||||
NETIF_F_SG | NETIF_F_GSO |
|
||||
NETIF_F_GSO_SOFTWARE | NETIF_F_HIGHDMA };
|
||||
const int overhead = MESSAGE_MINIMUM_LENGTH + sizeof(struct udphdr) +
|
||||
max(sizeof(struct ipv6hdr), sizeof(struct iphdr));
|
||||
|
||||
dev->netdev_ops = &netdev_ops;
|
||||
dev->header_ops = &ip_tunnel_header_ops;
|
||||
dev->hard_header_len = 0;
|
||||
dev->addr_len = 0;
|
||||
dev->needed_headroom = DATA_PACKET_HEAD_ROOM;
|
||||
dev->needed_tailroom = noise_encrypted_len(MESSAGE_PADDING_MULTIPLE);
|
||||
dev->type = ARPHRD_NONE;
|
||||
dev->flags = IFF_POINTOPOINT | IFF_NOARP;
|
||||
dev->priv_flags |= IFF_NO_QUEUE;
|
||||
dev->features |= NETIF_F_LLTX;
|
||||
dev->features |= WG_NETDEV_FEATURES;
|
||||
dev->hw_features |= WG_NETDEV_FEATURES;
|
||||
dev->hw_enc_features |= WG_NETDEV_FEATURES;
|
||||
dev->mtu = ETH_DATA_LEN - overhead;
|
||||
dev->max_mtu = round_down(INT_MAX, MESSAGE_PADDING_MULTIPLE) - overhead;
|
||||
|
||||
SET_NETDEV_DEVTYPE(dev, &device_type);
|
||||
|
||||
/* We need to keep the dst around in case of icmp replies. */
|
||||
netif_keep_dst(dev);
|
||||
|
||||
memset(wg, 0, sizeof(*wg));
|
||||
wg->dev = dev;
|
||||
}
|
||||
|
||||
static int wg_newlink(struct net *src_net, struct net_device *dev,
|
||||
struct nlattr *tb[], struct nlattr *data[],
|
||||
struct netlink_ext_ack *extack)
|
||||
{
|
||||
struct wg_device *wg = netdev_priv(dev);
|
||||
int ret = -ENOMEM;
|
||||
|
||||
rcu_assign_pointer(wg->creating_net, src_net);
|
||||
init_rwsem(&wg->static_identity.lock);
|
||||
mutex_init(&wg->socket_update_lock);
|
||||
mutex_init(&wg->device_update_lock);
|
||||
wg_allowedips_init(&wg->peer_allowedips);
|
||||
wg_cookie_checker_init(&wg->cookie_checker, wg);
|
||||
INIT_LIST_HEAD(&wg->peer_list);
|
||||
wg->device_update_gen = 1;
|
||||
|
||||
wg->peer_hashtable = wg_pubkey_hashtable_alloc();
|
||||
if (!wg->peer_hashtable)
|
||||
return ret;
|
||||
|
||||
wg->index_hashtable = wg_index_hashtable_alloc();
|
||||
if (!wg->index_hashtable)
|
||||
goto err_free_peer_hashtable;
|
||||
|
||||
dev->tstats = netdev_alloc_pcpu_stats(struct pcpu_sw_netstats);
|
||||
if (!dev->tstats)
|
||||
goto err_free_index_hashtable;
|
||||
|
||||
wg->handshake_receive_wq = alloc_workqueue("wg-kex-%s",
|
||||
WQ_CPU_INTENSIVE | WQ_FREEZABLE, 0, dev->name);
|
||||
if (!wg->handshake_receive_wq)
|
||||
goto err_free_tstats;
|
||||
|
||||
wg->handshake_send_wq = alloc_workqueue("wg-kex-%s",
|
||||
WQ_UNBOUND | WQ_FREEZABLE, 0, dev->name);
|
||||
if (!wg->handshake_send_wq)
|
||||
goto err_destroy_handshake_receive;
|
||||
|
||||
wg->packet_crypt_wq = alloc_workqueue("wg-crypt-%s",
|
||||
WQ_CPU_INTENSIVE | WQ_MEM_RECLAIM, 0, dev->name);
|
||||
if (!wg->packet_crypt_wq)
|
||||
goto err_destroy_handshake_send;
|
||||
|
||||
ret = wg_packet_queue_init(&wg->encrypt_queue, wg_packet_encrypt_worker,
|
||||
MAX_QUEUED_PACKETS);
|
||||
if (ret < 0)
|
||||
goto err_destroy_packet_crypt;
|
||||
|
||||
ret = wg_packet_queue_init(&wg->decrypt_queue, wg_packet_decrypt_worker,
|
||||
MAX_QUEUED_PACKETS);
|
||||
if (ret < 0)
|
||||
goto err_free_encrypt_queue;
|
||||
|
||||
ret = wg_packet_queue_init(&wg->handshake_queue, wg_packet_handshake_receive_worker,
|
||||
MAX_QUEUED_INCOMING_HANDSHAKES);
|
||||
if (ret < 0)
|
||||
goto err_free_decrypt_queue;
|
||||
|
||||
ret = wg_ratelimiter_init();
|
||||
if (ret < 0)
|
||||
goto err_free_handshake_queue;
|
||||
|
||||
ret = register_netdevice(dev);
|
||||
if (ret < 0)
|
||||
goto err_uninit_ratelimiter;
|
||||
|
||||
list_add(&wg->device_list, &device_list);
|
||||
|
||||
/* We wait until the end to assign priv_destructor, so that
|
||||
* register_netdevice doesn't call it for us if it fails.
|
||||
*/
|
||||
dev->priv_destructor = wg_destruct;
|
||||
|
||||
pr_debug("%s: Interface created\n", dev->name);
|
||||
return ret;
|
||||
|
||||
err_uninit_ratelimiter:
|
||||
wg_ratelimiter_uninit();
|
||||
err_free_handshake_queue:
|
||||
wg_packet_queue_free(&wg->handshake_queue, false);
|
||||
err_free_decrypt_queue:
|
||||
wg_packet_queue_free(&wg->decrypt_queue, false);
|
||||
err_free_encrypt_queue:
|
||||
wg_packet_queue_free(&wg->encrypt_queue, false);
|
||||
err_destroy_packet_crypt:
|
||||
destroy_workqueue(wg->packet_crypt_wq);
|
||||
err_destroy_handshake_send:
|
||||
destroy_workqueue(wg->handshake_send_wq);
|
||||
err_destroy_handshake_receive:
|
||||
destroy_workqueue(wg->handshake_receive_wq);
|
||||
err_free_tstats:
|
||||
free_percpu(dev->tstats);
|
||||
err_free_index_hashtable:
|
||||
kvfree(wg->index_hashtable);
|
||||
err_free_peer_hashtable:
|
||||
kvfree(wg->peer_hashtable);
|
||||
return ret;
|
||||
}
|
||||
|
||||
static struct rtnl_link_ops link_ops __read_mostly = {
|
||||
.kind = KBUILD_MODNAME,
|
||||
.priv_size = sizeof(struct wg_device),
|
||||
.setup = wg_setup,
|
||||
.newlink = wg_newlink,
|
||||
};
|
||||
|
||||
static void wg_netns_pre_exit(struct net *net)
|
||||
{
|
||||
struct wg_device *wg;
|
||||
struct wg_peer *peer;
|
||||
|
||||
rtnl_lock();
|
||||
list_for_each_entry(wg, &device_list, device_list) {
|
||||
if (rcu_access_pointer(wg->creating_net) == net) {
|
||||
pr_debug("%s: Creating namespace exiting\n", wg->dev->name);
|
||||
netif_carrier_off(wg->dev);
|
||||
mutex_lock(&wg->device_update_lock);
|
||||
rcu_assign_pointer(wg->creating_net, NULL);
|
||||
wg_socket_reinit(wg, NULL, NULL);
|
||||
list_for_each_entry(peer, &wg->peer_list, peer_list)
|
||||
wg_socket_clear_peer_endpoint_src(peer);
|
||||
mutex_unlock(&wg->device_update_lock);
|
||||
}
|
||||
}
|
||||
rtnl_unlock();
|
||||
}
|
||||
|
||||
static struct pernet_operations pernet_ops = {
|
||||
.pre_exit = wg_netns_pre_exit
|
||||
};
|
||||
|
||||
int __init wg_device_init(void)
|
||||
{
|
||||
int ret;
|
||||
|
||||
#ifdef CONFIG_PM_SLEEP
|
||||
ret = register_pm_notifier(&pm_notifier);
|
||||
if (ret)
|
||||
return ret;
|
||||
#endif
|
||||
|
||||
ret = register_pernet_device(&pernet_ops);
|
||||
if (ret)
|
||||
goto error_pm;
|
||||
|
||||
ret = rtnl_link_register(&link_ops);
|
||||
if (ret)
|
||||
goto error_pernet;
|
||||
|
||||
return 0;
|
||||
|
||||
error_pernet:
|
||||
unregister_pernet_device(&pernet_ops);
|
||||
error_pm:
|
||||
#ifdef CONFIG_PM_SLEEP
|
||||
unregister_pm_notifier(&pm_notifier);
|
||||
#endif
|
||||
return ret;
|
||||
}
|
||||
|
||||
void wg_device_uninit(void)
|
||||
{
|
||||
rtnl_link_unregister(&link_ops);
|
||||
unregister_pernet_device(&pernet_ops);
|
||||
#ifdef CONFIG_PM_SLEEP
|
||||
unregister_pm_notifier(&pm_notifier);
|
||||
#endif
|
||||
rcu_barrier();
|
||||
}
|
||||
62
drivers/net/wireguard/device.h
Normal file
62
drivers/net/wireguard/device.h
Normal file
|
|
@ -0,0 +1,62 @@
|
|||
/* SPDX-License-Identifier: GPL-2.0 */
|
||||
/*
|
||||
* Copyright (C) 2015-2019 Jason A. Donenfeld <Jason@zx2c4.com>. All Rights Reserved.
|
||||
*/
|
||||
|
||||
#ifndef _WG_DEVICE_H
|
||||
#define _WG_DEVICE_H
|
||||
|
||||
#include "noise.h"
|
||||
#include "allowedips.h"
|
||||
#include "peerlookup.h"
|
||||
#include "cookie.h"
|
||||
|
||||
#include <linux/types.h>
|
||||
#include <linux/netdevice.h>
|
||||
#include <linux/workqueue.h>
|
||||
#include <linux/mutex.h>
|
||||
#include <linux/net.h>
|
||||
#include <linux/ptr_ring.h>
|
||||
|
||||
struct wg_device;
|
||||
|
||||
struct multicore_worker {
|
||||
void *ptr;
|
||||
struct work_struct work;
|
||||
};
|
||||
|
||||
struct crypt_queue {
|
||||
struct ptr_ring ring;
|
||||
struct multicore_worker __percpu *worker;
|
||||
int last_cpu;
|
||||
};
|
||||
|
||||
struct prev_queue {
|
||||
struct sk_buff *head, *tail, *peeked;
|
||||
struct { struct sk_buff *next, *prev; } empty; // Match first 2 members of struct sk_buff.
|
||||
atomic_t count;
|
||||
};
|
||||
|
||||
struct wg_device {
|
||||
struct net_device *dev;
|
||||
struct crypt_queue encrypt_queue, decrypt_queue, handshake_queue;
|
||||
struct sock __rcu *sock4, *sock6;
|
||||
struct net __rcu *creating_net;
|
||||
struct noise_static_identity static_identity;
|
||||
struct workqueue_struct *packet_crypt_wq,*handshake_receive_wq, *handshake_send_wq;
|
||||
struct cookie_checker cookie_checker;
|
||||
struct pubkey_hashtable *peer_hashtable;
|
||||
struct index_hashtable *index_hashtable;
|
||||
struct allowedips peer_allowedips;
|
||||
struct mutex device_update_lock, socket_update_lock;
|
||||
struct list_head device_list, peer_list;
|
||||
atomic_t handshake_queue_len;
|
||||
unsigned int num_peers, device_update_gen;
|
||||
u32 fwmark;
|
||||
u16 incoming_port;
|
||||
};
|
||||
|
||||
int wg_device_init(void);
|
||||
void wg_device_uninit(void);
|
||||
|
||||
#endif /* _WG_DEVICE_H */
|
||||
78
drivers/net/wireguard/main.c
Normal file
78
drivers/net/wireguard/main.c
Normal file
|
|
@ -0,0 +1,78 @@
|
|||
// SPDX-License-Identifier: GPL-2.0
|
||||
/*
|
||||
* Copyright (C) 2015-2019 Jason A. Donenfeld <Jason@zx2c4.com>. All Rights Reserved.
|
||||
*/
|
||||
|
||||
#include "version.h"
|
||||
#include "device.h"
|
||||
#include "noise.h"
|
||||
#include "queueing.h"
|
||||
#include "ratelimiter.h"
|
||||
#include "netlink.h"
|
||||
|
||||
#include <uapi/linux/wireguard.h>
|
||||
|
||||
#include <linux/init.h>
|
||||
#include <linux/module.h>
|
||||
#include <linux/genetlink.h>
|
||||
#include <net/rtnetlink.h>
|
||||
|
||||
static int __init mod_init(void)
|
||||
{
|
||||
int ret;
|
||||
|
||||
ret = wg_allowedips_slab_init();
|
||||
if (ret < 0)
|
||||
goto err_allowedips;
|
||||
|
||||
#ifdef DEBUG
|
||||
ret = -ENOTRECOVERABLE;
|
||||
if (!wg_allowedips_selftest() || !wg_packet_counter_selftest() ||
|
||||
!wg_ratelimiter_selftest())
|
||||
goto err_peer;
|
||||
#endif
|
||||
wg_noise_init();
|
||||
|
||||
ret = wg_peer_init();
|
||||
if (ret < 0)
|
||||
goto err_peer;
|
||||
|
||||
ret = wg_device_init();
|
||||
if (ret < 0)
|
||||
goto err_device;
|
||||
|
||||
ret = wg_genetlink_init();
|
||||
if (ret < 0)
|
||||
goto err_netlink;
|
||||
|
||||
pr_info("WireGuard " WIREGUARD_VERSION " loaded. See www.wireguard.com for information.\n");
|
||||
pr_info("Copyright (C) 2015-2019 Jason A. Donenfeld <Jason@zx2c4.com>. All Rights Reserved.\n");
|
||||
|
||||
return 0;
|
||||
|
||||
err_netlink:
|
||||
wg_device_uninit();
|
||||
err_device:
|
||||
wg_peer_uninit();
|
||||
err_peer:
|
||||
wg_allowedips_slab_uninit();
|
||||
err_allowedips:
|
||||
return ret;
|
||||
}
|
||||
|
||||
static void __exit mod_exit(void)
|
||||
{
|
||||
wg_genetlink_uninit();
|
||||
wg_device_uninit();
|
||||
wg_peer_uninit();
|
||||
wg_allowedips_slab_uninit();
|
||||
}
|
||||
|
||||
module_init(mod_init);
|
||||
module_exit(mod_exit);
|
||||
MODULE_LICENSE("GPL v2");
|
||||
MODULE_DESCRIPTION("WireGuard secure network tunnel");
|
||||
MODULE_AUTHOR("Jason A. Donenfeld <Jason@zx2c4.com>");
|
||||
MODULE_VERSION(WIREGUARD_VERSION);
|
||||
MODULE_ALIAS_RTNL_LINK(KBUILD_MODNAME);
|
||||
MODULE_ALIAS_GENL_FAMILY(WG_GENL_NAME);
|
||||
128
drivers/net/wireguard/messages.h
Normal file
128
drivers/net/wireguard/messages.h
Normal file
|
|
@ -0,0 +1,128 @@
|
|||
/* SPDX-License-Identifier: GPL-2.0 */
|
||||
/*
|
||||
* Copyright (C) 2015-2019 Jason A. Donenfeld <Jason@zx2c4.com>. All Rights Reserved.
|
||||
*/
|
||||
|
||||
#ifndef _WG_MESSAGES_H
|
||||
#define _WG_MESSAGES_H
|
||||
|
||||
#include <crypto/curve25519.h>
|
||||
#include <crypto/chacha20poly1305.h>
|
||||
#include <crypto/blake2s.h>
|
||||
|
||||
#include <linux/kernel.h>
|
||||
#include <linux/param.h>
|
||||
#include <linux/skbuff.h>
|
||||
|
||||
enum noise_lengths {
|
||||
NOISE_PUBLIC_KEY_LEN = CURVE25519_KEY_SIZE,
|
||||
NOISE_SYMMETRIC_KEY_LEN = CHACHA20POLY1305_KEY_SIZE,
|
||||
NOISE_TIMESTAMP_LEN = sizeof(u64) + sizeof(u32),
|
||||
NOISE_AUTHTAG_LEN = CHACHA20POLY1305_AUTHTAG_SIZE,
|
||||
NOISE_HASH_LEN = BLAKE2S_HASH_SIZE
|
||||
};
|
||||
|
||||
#define noise_encrypted_len(plain_len) ((plain_len) + NOISE_AUTHTAG_LEN)
|
||||
|
||||
enum cookie_values {
|
||||
COOKIE_SECRET_MAX_AGE = 2 * 60,
|
||||
COOKIE_SECRET_LATENCY = 5,
|
||||
COOKIE_NONCE_LEN = XCHACHA20POLY1305_NONCE_SIZE,
|
||||
COOKIE_LEN = 16
|
||||
};
|
||||
|
||||
enum counter_values {
|
||||
COUNTER_BITS_TOTAL = 8192,
|
||||
COUNTER_REDUNDANT_BITS = BITS_PER_LONG,
|
||||
COUNTER_WINDOW_SIZE = COUNTER_BITS_TOTAL - COUNTER_REDUNDANT_BITS
|
||||
};
|
||||
|
||||
enum limits {
|
||||
REKEY_AFTER_MESSAGES = 1ULL << 60,
|
||||
REJECT_AFTER_MESSAGES = U64_MAX - COUNTER_WINDOW_SIZE - 1,
|
||||
REKEY_TIMEOUT = 5,
|
||||
REKEY_TIMEOUT_JITTER_MAX_JIFFIES = HZ / 3,
|
||||
REKEY_AFTER_TIME = 120,
|
||||
REJECT_AFTER_TIME = 180,
|
||||
INITIATIONS_PER_SECOND = 50,
|
||||
MAX_PEERS_PER_DEVICE = 1U << 20,
|
||||
KEEPALIVE_TIMEOUT = 10,
|
||||
MAX_TIMER_HANDSHAKES = 90 / REKEY_TIMEOUT,
|
||||
MAX_QUEUED_INCOMING_HANDSHAKES = 4096, /* TODO: replace this with DQL */
|
||||
MAX_STAGED_PACKETS = 128,
|
||||
MAX_QUEUED_PACKETS = 1024 /* TODO: replace this with DQL */
|
||||
};
|
||||
|
||||
enum message_type {
|
||||
MESSAGE_INVALID = 0,
|
||||
MESSAGE_HANDSHAKE_INITIATION = 1,
|
||||
MESSAGE_HANDSHAKE_RESPONSE = 2,
|
||||
MESSAGE_HANDSHAKE_COOKIE = 3,
|
||||
MESSAGE_DATA = 4
|
||||
};
|
||||
|
||||
struct message_header {
|
||||
/* The actual layout of this that we want is:
|
||||
* u8 type
|
||||
* u8 reserved_zero[3]
|
||||
*
|
||||
* But it turns out that by encoding this as little endian,
|
||||
* we achieve the same thing, and it makes checking faster.
|
||||
*/
|
||||
__le32 type;
|
||||
};
|
||||
|
||||
struct message_macs {
|
||||
u8 mac1[COOKIE_LEN];
|
||||
u8 mac2[COOKIE_LEN];
|
||||
};
|
||||
|
||||
struct message_handshake_initiation {
|
||||
struct message_header header;
|
||||
__le32 sender_index;
|
||||
u8 unencrypted_ephemeral[NOISE_PUBLIC_KEY_LEN];
|
||||
u8 encrypted_static[noise_encrypted_len(NOISE_PUBLIC_KEY_LEN)];
|
||||
u8 encrypted_timestamp[noise_encrypted_len(NOISE_TIMESTAMP_LEN)];
|
||||
struct message_macs macs;
|
||||
};
|
||||
|
||||
struct message_handshake_response {
|
||||
struct message_header header;
|
||||
__le32 sender_index;
|
||||
__le32 receiver_index;
|
||||
u8 unencrypted_ephemeral[NOISE_PUBLIC_KEY_LEN];
|
||||
u8 encrypted_nothing[noise_encrypted_len(0)];
|
||||
struct message_macs macs;
|
||||
};
|
||||
|
||||
struct message_handshake_cookie {
|
||||
struct message_header header;
|
||||
__le32 receiver_index;
|
||||
u8 nonce[COOKIE_NONCE_LEN];
|
||||
u8 encrypted_cookie[noise_encrypted_len(COOKIE_LEN)];
|
||||
};
|
||||
|
||||
struct message_data {
|
||||
struct message_header header;
|
||||
__le32 key_idx;
|
||||
__le64 counter;
|
||||
u8 encrypted_data[];
|
||||
};
|
||||
|
||||
#define message_data_len(plain_len) \
|
||||
(noise_encrypted_len(plain_len) + sizeof(struct message_data))
|
||||
|
||||
enum message_alignments {
|
||||
MESSAGE_PADDING_MULTIPLE = 16,
|
||||
MESSAGE_MINIMUM_LENGTH = message_data_len(0)
|
||||
};
|
||||
|
||||
#define SKB_HEADER_LEN \
|
||||
(max(sizeof(struct iphdr), sizeof(struct ipv6hdr)) + \
|
||||
sizeof(struct udphdr) + NET_SKB_PAD)
|
||||
#define DATA_PACKET_HEAD_ROOM \
|
||||
ALIGN(sizeof(struct message_data) + SKB_HEADER_LEN, 4)
|
||||
|
||||
enum { HANDSHAKE_DSCP = 0x88 /* AF41, plus 00 ECN */ };
|
||||
|
||||
#endif /* _WG_MESSAGES_H */
|
||||
649
drivers/net/wireguard/netlink.c
Normal file
649
drivers/net/wireguard/netlink.c
Normal file
|
|
@ -0,0 +1,649 @@
|
|||
// SPDX-License-Identifier: GPL-2.0
|
||||
/*
|
||||
* Copyright (C) 2015-2019 Jason A. Donenfeld <Jason@zx2c4.com>. All Rights Reserved.
|
||||
*/
|
||||
|
||||
#include "netlink.h"
|
||||
#include "device.h"
|
||||
#include "peer.h"
|
||||
#include "socket.h"
|
||||
#include "queueing.h"
|
||||
#include "messages.h"
|
||||
|
||||
#include <uapi/linux/wireguard.h>
|
||||
|
||||
#include <linux/if.h>
|
||||
#include <net/genetlink.h>
|
||||
#include <net/sock.h>
|
||||
#include <crypto/algapi.h>
|
||||
|
||||
static struct genl_family genl_family;
|
||||
|
||||
static const struct nla_policy device_policy[WGDEVICE_A_MAX + 1] = {
|
||||
[WGDEVICE_A_IFINDEX] = { .type = NLA_U32 },
|
||||
[WGDEVICE_A_IFNAME] = { .type = NLA_NUL_STRING, .len = IFNAMSIZ - 1 },
|
||||
[WGDEVICE_A_PRIVATE_KEY] = NLA_POLICY_EXACT_LEN(NOISE_PUBLIC_KEY_LEN),
|
||||
[WGDEVICE_A_PUBLIC_KEY] = NLA_POLICY_EXACT_LEN(NOISE_PUBLIC_KEY_LEN),
|
||||
[WGDEVICE_A_FLAGS] = { .type = NLA_U32 },
|
||||
[WGDEVICE_A_LISTEN_PORT] = { .type = NLA_U16 },
|
||||
[WGDEVICE_A_FWMARK] = { .type = NLA_U32 },
|
||||
[WGDEVICE_A_PEERS] = { .type = NLA_NESTED }
|
||||
};
|
||||
|
||||
static const struct nla_policy peer_policy[WGPEER_A_MAX + 1] = {
|
||||
[WGPEER_A_PUBLIC_KEY] = NLA_POLICY_EXACT_LEN(NOISE_PUBLIC_KEY_LEN),
|
||||
[WGPEER_A_PRESHARED_KEY] = NLA_POLICY_EXACT_LEN(NOISE_SYMMETRIC_KEY_LEN),
|
||||
[WGPEER_A_FLAGS] = { .type = NLA_U32 },
|
||||
[WGPEER_A_ENDPOINT] = NLA_POLICY_MIN_LEN(sizeof(struct sockaddr)),
|
||||
[WGPEER_A_PERSISTENT_KEEPALIVE_INTERVAL] = { .type = NLA_U16 },
|
||||
[WGPEER_A_LAST_HANDSHAKE_TIME] = NLA_POLICY_EXACT_LEN(sizeof(struct __kernel_timespec)),
|
||||
[WGPEER_A_RX_BYTES] = { .type = NLA_U64 },
|
||||
[WGPEER_A_TX_BYTES] = { .type = NLA_U64 },
|
||||
[WGPEER_A_ALLOWEDIPS] = { .type = NLA_NESTED },
|
||||
[WGPEER_A_PROTOCOL_VERSION] = { .type = NLA_U32 }
|
||||
};
|
||||
|
||||
static const struct nla_policy allowedip_policy[WGALLOWEDIP_A_MAX + 1] = {
|
||||
[WGALLOWEDIP_A_FAMILY] = { .type = NLA_U16 },
|
||||
[WGALLOWEDIP_A_IPADDR] = NLA_POLICY_MIN_LEN(sizeof(struct in_addr)),
|
||||
[WGALLOWEDIP_A_CIDR_MASK] = { .type = NLA_U8 }
|
||||
};
|
||||
|
||||
static struct wg_device *lookup_interface(struct nlattr **attrs,
|
||||
struct sk_buff *skb)
|
||||
{
|
||||
struct net_device *dev = NULL;
|
||||
|
||||
if (!attrs[WGDEVICE_A_IFINDEX] == !attrs[WGDEVICE_A_IFNAME])
|
||||
return ERR_PTR(-EBADR);
|
||||
if (attrs[WGDEVICE_A_IFINDEX])
|
||||
dev = dev_get_by_index(sock_net(skb->sk),
|
||||
nla_get_u32(attrs[WGDEVICE_A_IFINDEX]));
|
||||
else if (attrs[WGDEVICE_A_IFNAME])
|
||||
dev = dev_get_by_name(sock_net(skb->sk),
|
||||
nla_data(attrs[WGDEVICE_A_IFNAME]));
|
||||
if (!dev)
|
||||
return ERR_PTR(-ENODEV);
|
||||
if (!dev->rtnl_link_ops || !dev->rtnl_link_ops->kind ||
|
||||
strcmp(dev->rtnl_link_ops->kind, KBUILD_MODNAME)) {
|
||||
dev_put(dev);
|
||||
return ERR_PTR(-EOPNOTSUPP);
|
||||
}
|
||||
return netdev_priv(dev);
|
||||
}
|
||||
|
||||
static int get_allowedips(struct sk_buff *skb, const u8 *ip, u8 cidr,
|
||||
int family)
|
||||
{
|
||||
struct nlattr *allowedip_nest;
|
||||
|
||||
allowedip_nest = nla_nest_start(skb, 0);
|
||||
if (!allowedip_nest)
|
||||
return -EMSGSIZE;
|
||||
|
||||
if (nla_put_u8(skb, WGALLOWEDIP_A_CIDR_MASK, cidr) ||
|
||||
nla_put_u16(skb, WGALLOWEDIP_A_FAMILY, family) ||
|
||||
nla_put(skb, WGALLOWEDIP_A_IPADDR, family == AF_INET6 ?
|
||||
sizeof(struct in6_addr) : sizeof(struct in_addr), ip)) {
|
||||
nla_nest_cancel(skb, allowedip_nest);
|
||||
return -EMSGSIZE;
|
||||
}
|
||||
|
||||
nla_nest_end(skb, allowedip_nest);
|
||||
return 0;
|
||||
}
|
||||
|
||||
struct dump_ctx {
|
||||
struct wg_device *wg;
|
||||
struct wg_peer *next_peer;
|
||||
u64 allowedips_seq;
|
||||
struct allowedips_node *next_allowedip;
|
||||
};
|
||||
|
||||
#define DUMP_CTX(cb) ((struct dump_ctx *)(cb)->args)
|
||||
|
||||
static int
|
||||
get_peer(struct wg_peer *peer, struct sk_buff *skb, struct dump_ctx *ctx)
|
||||
{
|
||||
|
||||
struct nlattr *allowedips_nest, *peer_nest = nla_nest_start(skb, 0);
|
||||
struct allowedips_node *allowedips_node = ctx->next_allowedip;
|
||||
bool fail;
|
||||
|
||||
if (!peer_nest)
|
||||
return -EMSGSIZE;
|
||||
|
||||
down_read(&peer->handshake.lock);
|
||||
fail = nla_put(skb, WGPEER_A_PUBLIC_KEY, NOISE_PUBLIC_KEY_LEN,
|
||||
peer->handshake.remote_static);
|
||||
up_read(&peer->handshake.lock);
|
||||
if (fail)
|
||||
goto err;
|
||||
|
||||
if (!allowedips_node) {
|
||||
const struct __kernel_timespec last_handshake = {
|
||||
.tv_sec = peer->walltime_last_handshake.tv_sec,
|
||||
.tv_nsec = peer->walltime_last_handshake.tv_nsec
|
||||
};
|
||||
|
||||
down_read(&peer->handshake.lock);
|
||||
fail = nla_put(skb, WGPEER_A_PRESHARED_KEY,
|
||||
NOISE_SYMMETRIC_KEY_LEN,
|
||||
peer->handshake.preshared_key);
|
||||
up_read(&peer->handshake.lock);
|
||||
if (fail)
|
||||
goto err;
|
||||
|
||||
if (nla_put(skb, WGPEER_A_LAST_HANDSHAKE_TIME,
|
||||
sizeof(last_handshake), &last_handshake) ||
|
||||
nla_put_u16(skb, WGPEER_A_PERSISTENT_KEEPALIVE_INTERVAL,
|
||||
peer->persistent_keepalive_interval) ||
|
||||
nla_put_u64_64bit(skb, WGPEER_A_TX_BYTES, peer->tx_bytes,
|
||||
WGPEER_A_UNSPEC) ||
|
||||
nla_put_u64_64bit(skb, WGPEER_A_RX_BYTES, peer->rx_bytes,
|
||||
WGPEER_A_UNSPEC) ||
|
||||
nla_put_u32(skb, WGPEER_A_PROTOCOL_VERSION, 1))
|
||||
goto err;
|
||||
|
||||
read_lock_bh(&peer->endpoint_lock);
|
||||
if (peer->endpoint.addr.sa_family == AF_INET)
|
||||
fail = nla_put(skb, WGPEER_A_ENDPOINT,
|
||||
sizeof(peer->endpoint.addr4),
|
||||
&peer->endpoint.addr4);
|
||||
else if (peer->endpoint.addr.sa_family == AF_INET6)
|
||||
fail = nla_put(skb, WGPEER_A_ENDPOINT,
|
||||
sizeof(peer->endpoint.addr6),
|
||||
&peer->endpoint.addr6);
|
||||
read_unlock_bh(&peer->endpoint_lock);
|
||||
if (fail)
|
||||
goto err;
|
||||
allowedips_node =
|
||||
list_first_entry_or_null(&peer->allowedips_list,
|
||||
struct allowedips_node, peer_list);
|
||||
}
|
||||
if (!allowedips_node)
|
||||
goto no_allowedips;
|
||||
if (!ctx->allowedips_seq)
|
||||
ctx->allowedips_seq = ctx->wg->peer_allowedips.seq;
|
||||
else if (ctx->allowedips_seq != ctx->wg->peer_allowedips.seq)
|
||||
goto no_allowedips;
|
||||
|
||||
allowedips_nest = nla_nest_start(skb, WGPEER_A_ALLOWEDIPS);
|
||||
if (!allowedips_nest)
|
||||
goto err;
|
||||
|
||||
list_for_each_entry_from(allowedips_node, &peer->allowedips_list,
|
||||
peer_list) {
|
||||
u8 cidr, ip[16] __aligned(__alignof(u64));
|
||||
int family;
|
||||
|
||||
family = wg_allowedips_read_node(allowedips_node, ip, &cidr);
|
||||
if (get_allowedips(skb, ip, cidr, family)) {
|
||||
nla_nest_end(skb, allowedips_nest);
|
||||
nla_nest_end(skb, peer_nest);
|
||||
ctx->next_allowedip = allowedips_node;
|
||||
return -EMSGSIZE;
|
||||
}
|
||||
}
|
||||
nla_nest_end(skb, allowedips_nest);
|
||||
no_allowedips:
|
||||
nla_nest_end(skb, peer_nest);
|
||||
ctx->next_allowedip = NULL;
|
||||
ctx->allowedips_seq = 0;
|
||||
return 0;
|
||||
err:
|
||||
nla_nest_cancel(skb, peer_nest);
|
||||
return -EMSGSIZE;
|
||||
}
|
||||
|
||||
static int wg_get_device_start(struct netlink_callback *cb)
|
||||
{
|
||||
struct nlattr **attrs = genl_family_attrbuf(&genl_family);
|
||||
struct wg_device *wg;
|
||||
int ret;
|
||||
|
||||
ret = nlmsg_parse(cb->nlh, GENL_HDRLEN + genl_family.hdrsize, attrs,
|
||||
genl_family.maxattr, device_policy, NULL);
|
||||
if (ret < 0)
|
||||
return ret;
|
||||
wg = lookup_interface(attrs, cb->skb);
|
||||
if (IS_ERR(wg))
|
||||
return PTR_ERR(wg);
|
||||
DUMP_CTX(cb)->wg = wg;
|
||||
return 0;
|
||||
}
|
||||
|
||||
static int wg_get_device_dump(struct sk_buff *skb, struct netlink_callback *cb)
|
||||
{
|
||||
struct wg_peer *peer, *next_peer_cursor;
|
||||
struct dump_ctx *ctx = DUMP_CTX(cb);
|
||||
struct wg_device *wg = ctx->wg;
|
||||
struct nlattr *peers_nest;
|
||||
int ret = -EMSGSIZE;
|
||||
bool done = true;
|
||||
void *hdr;
|
||||
|
||||
rtnl_lock();
|
||||
mutex_lock(&wg->device_update_lock);
|
||||
cb->seq = wg->device_update_gen;
|
||||
next_peer_cursor = ctx->next_peer;
|
||||
|
||||
hdr = genlmsg_put(skb, NETLINK_CB(cb->skb).portid, cb->nlh->nlmsg_seq,
|
||||
&genl_family, NLM_F_MULTI, WG_CMD_GET_DEVICE);
|
||||
if (!hdr)
|
||||
goto out;
|
||||
genl_dump_check_consistent(cb, hdr);
|
||||
|
||||
if (!ctx->next_peer) {
|
||||
if (nla_put_u16(skb, WGDEVICE_A_LISTEN_PORT,
|
||||
wg->incoming_port) ||
|
||||
nla_put_u32(skb, WGDEVICE_A_FWMARK, wg->fwmark) ||
|
||||
nla_put_u32(skb, WGDEVICE_A_IFINDEX, wg->dev->ifindex) ||
|
||||
nla_put_string(skb, WGDEVICE_A_IFNAME, wg->dev->name))
|
||||
goto out;
|
||||
|
||||
down_read(&wg->static_identity.lock);
|
||||
if (wg->static_identity.has_identity) {
|
||||
if (nla_put(skb, WGDEVICE_A_PRIVATE_KEY,
|
||||
NOISE_PUBLIC_KEY_LEN,
|
||||
wg->static_identity.static_private) ||
|
||||
nla_put(skb, WGDEVICE_A_PUBLIC_KEY,
|
||||
NOISE_PUBLIC_KEY_LEN,
|
||||
wg->static_identity.static_public)) {
|
||||
up_read(&wg->static_identity.lock);
|
||||
goto out;
|
||||
}
|
||||
}
|
||||
up_read(&wg->static_identity.lock);
|
||||
}
|
||||
|
||||
peers_nest = nla_nest_start(skb, WGDEVICE_A_PEERS);
|
||||
if (!peers_nest)
|
||||
goto out;
|
||||
ret = 0;
|
||||
lockdep_assert_held(&wg->device_update_lock);
|
||||
/* If the last cursor was removed in peer_remove or peer_remove_all, then
|
||||
* we just treat this the same as there being no more peers left. The
|
||||
* reason is that seq_nr should indicate to userspace that this isn't a
|
||||
* coherent dump anyway, so they'll try again.
|
||||
*/
|
||||
if (list_empty(&wg->peer_list) ||
|
||||
(ctx->next_peer && ctx->next_peer->is_dead)) {
|
||||
nla_nest_cancel(skb, peers_nest);
|
||||
goto out;
|
||||
}
|
||||
peer = list_prepare_entry(ctx->next_peer, &wg->peer_list, peer_list);
|
||||
list_for_each_entry_continue(peer, &wg->peer_list, peer_list) {
|
||||
if (get_peer(peer, skb, ctx)) {
|
||||
done = false;
|
||||
break;
|
||||
}
|
||||
next_peer_cursor = peer;
|
||||
}
|
||||
nla_nest_end(skb, peers_nest);
|
||||
|
||||
out:
|
||||
if (!ret && !done && next_peer_cursor)
|
||||
wg_peer_get(next_peer_cursor);
|
||||
wg_peer_put(ctx->next_peer);
|
||||
mutex_unlock(&wg->device_update_lock);
|
||||
rtnl_unlock();
|
||||
|
||||
if (ret) {
|
||||
genlmsg_cancel(skb, hdr);
|
||||
return ret;
|
||||
}
|
||||
genlmsg_end(skb, hdr);
|
||||
if (done) {
|
||||
ctx->next_peer = NULL;
|
||||
return 0;
|
||||
}
|
||||
ctx->next_peer = next_peer_cursor;
|
||||
return skb->len;
|
||||
|
||||
/* At this point, we can't really deal ourselves with safely zeroing out
|
||||
* the private key material after usage. This will need an additional API
|
||||
* in the kernel for marking skbs as zero_on_free.
|
||||
*/
|
||||
}
|
||||
|
||||
static int wg_get_device_done(struct netlink_callback *cb)
|
||||
{
|
||||
struct dump_ctx *ctx = DUMP_CTX(cb);
|
||||
|
||||
if (ctx->wg)
|
||||
dev_put(ctx->wg->dev);
|
||||
wg_peer_put(ctx->next_peer);
|
||||
return 0;
|
||||
}
|
||||
|
||||
static int set_port(struct wg_device *wg, u16 port)
|
||||
{
|
||||
struct wg_peer *peer;
|
||||
|
||||
if (wg->incoming_port == port)
|
||||
return 0;
|
||||
list_for_each_entry(peer, &wg->peer_list, peer_list)
|
||||
wg_socket_clear_peer_endpoint_src(peer);
|
||||
if (!netif_running(wg->dev)) {
|
||||
wg->incoming_port = port;
|
||||
return 0;
|
||||
}
|
||||
return wg_socket_init(wg, port);
|
||||
}
|
||||
|
||||
static int set_allowedip(struct wg_peer *peer, struct nlattr **attrs)
|
||||
{
|
||||
int ret = -EINVAL;
|
||||
u16 family;
|
||||
u8 cidr;
|
||||
|
||||
if (!attrs[WGALLOWEDIP_A_FAMILY] || !attrs[WGALLOWEDIP_A_IPADDR] ||
|
||||
!attrs[WGALLOWEDIP_A_CIDR_MASK])
|
||||
return ret;
|
||||
family = nla_get_u16(attrs[WGALLOWEDIP_A_FAMILY]);
|
||||
cidr = nla_get_u8(attrs[WGALLOWEDIP_A_CIDR_MASK]);
|
||||
|
||||
if (family == AF_INET && cidr <= 32 &&
|
||||
nla_len(attrs[WGALLOWEDIP_A_IPADDR]) == sizeof(struct in_addr))
|
||||
ret = wg_allowedips_insert_v4(
|
||||
&peer->device->peer_allowedips,
|
||||
nla_data(attrs[WGALLOWEDIP_A_IPADDR]), cidr, peer,
|
||||
&peer->device->device_update_lock);
|
||||
else if (family == AF_INET6 && cidr <= 128 &&
|
||||
nla_len(attrs[WGALLOWEDIP_A_IPADDR]) == sizeof(struct in6_addr))
|
||||
ret = wg_allowedips_insert_v6(
|
||||
&peer->device->peer_allowedips,
|
||||
nla_data(attrs[WGALLOWEDIP_A_IPADDR]), cidr, peer,
|
||||
&peer->device->device_update_lock);
|
||||
|
||||
return ret;
|
||||
}
|
||||
|
||||
static int set_peer(struct wg_device *wg, struct nlattr **attrs)
|
||||
{
|
||||
u8 *public_key = NULL, *preshared_key = NULL;
|
||||
struct wg_peer *peer = NULL;
|
||||
u32 flags = 0;
|
||||
int ret;
|
||||
|
||||
ret = -EINVAL;
|
||||
if (attrs[WGPEER_A_PUBLIC_KEY] &&
|
||||
nla_len(attrs[WGPEER_A_PUBLIC_KEY]) == NOISE_PUBLIC_KEY_LEN)
|
||||
public_key = nla_data(attrs[WGPEER_A_PUBLIC_KEY]);
|
||||
else
|
||||
goto out;
|
||||
if (attrs[WGPEER_A_PRESHARED_KEY] &&
|
||||
nla_len(attrs[WGPEER_A_PRESHARED_KEY]) == NOISE_SYMMETRIC_KEY_LEN)
|
||||
preshared_key = nla_data(attrs[WGPEER_A_PRESHARED_KEY]);
|
||||
|
||||
if (attrs[WGPEER_A_FLAGS])
|
||||
flags = nla_get_u32(attrs[WGPEER_A_FLAGS]);
|
||||
ret = -EOPNOTSUPP;
|
||||
if (flags & ~__WGPEER_F_ALL)
|
||||
goto out;
|
||||
|
||||
ret = -EPFNOSUPPORT;
|
||||
if (attrs[WGPEER_A_PROTOCOL_VERSION]) {
|
||||
if (nla_get_u32(attrs[WGPEER_A_PROTOCOL_VERSION]) != 1)
|
||||
goto out;
|
||||
}
|
||||
|
||||
peer = wg_pubkey_hashtable_lookup(wg->peer_hashtable,
|
||||
nla_data(attrs[WGPEER_A_PUBLIC_KEY]));
|
||||
ret = 0;
|
||||
if (!peer) { /* Peer doesn't exist yet. Add a new one. */
|
||||
if (flags & (WGPEER_F_REMOVE_ME | WGPEER_F_UPDATE_ONLY))
|
||||
goto out;
|
||||
|
||||
/* The peer is new, so there aren't allowed IPs to remove. */
|
||||
flags &= ~WGPEER_F_REPLACE_ALLOWEDIPS;
|
||||
|
||||
down_read(&wg->static_identity.lock);
|
||||
if (wg->static_identity.has_identity &&
|
||||
!memcmp(nla_data(attrs[WGPEER_A_PUBLIC_KEY]),
|
||||
wg->static_identity.static_public,
|
||||
NOISE_PUBLIC_KEY_LEN)) {
|
||||
/* We silently ignore peers that have the same public
|
||||
* key as the device. The reason we do it silently is
|
||||
* that we'd like for people to be able to reuse the
|
||||
* same set of API calls across peers.
|
||||
*/
|
||||
up_read(&wg->static_identity.lock);
|
||||
ret = 0;
|
||||
goto out;
|
||||
}
|
||||
up_read(&wg->static_identity.lock);
|
||||
|
||||
peer = wg_peer_create(wg, public_key, preshared_key);
|
||||
if (IS_ERR(peer)) {
|
||||
ret = PTR_ERR(peer);
|
||||
peer = NULL;
|
||||
goto out;
|
||||
}
|
||||
/* Take additional reference, as though we've just been
|
||||
* looked up.
|
||||
*/
|
||||
wg_peer_get(peer);
|
||||
}
|
||||
|
||||
if (flags & WGPEER_F_REMOVE_ME) {
|
||||
wg_peer_remove(peer);
|
||||
goto out;
|
||||
}
|
||||
|
||||
if (preshared_key) {
|
||||
down_write(&peer->handshake.lock);
|
||||
memcpy(&peer->handshake.preshared_key, preshared_key,
|
||||
NOISE_SYMMETRIC_KEY_LEN);
|
||||
up_write(&peer->handshake.lock);
|
||||
}
|
||||
|
||||
if (attrs[WGPEER_A_ENDPOINT]) {
|
||||
struct sockaddr *addr = nla_data(attrs[WGPEER_A_ENDPOINT]);
|
||||
size_t len = nla_len(attrs[WGPEER_A_ENDPOINT]);
|
||||
struct endpoint endpoint = { { { 0 } } };
|
||||
|
||||
if (len == sizeof(struct sockaddr_in) && addr->sa_family == AF_INET) {
|
||||
endpoint.addr4 = *(struct sockaddr_in *)addr;
|
||||
wg_socket_set_peer_endpoint(peer, &endpoint);
|
||||
} else if (len == sizeof(struct sockaddr_in6) && addr->sa_family == AF_INET6) {
|
||||
endpoint.addr6 = *(struct sockaddr_in6 *)addr;
|
||||
wg_socket_set_peer_endpoint(peer, &endpoint);
|
||||
}
|
||||
}
|
||||
|
||||
if (flags & WGPEER_F_REPLACE_ALLOWEDIPS)
|
||||
wg_allowedips_remove_by_peer(&wg->peer_allowedips, peer,
|
||||
&wg->device_update_lock);
|
||||
|
||||
if (attrs[WGPEER_A_ALLOWEDIPS]) {
|
||||
struct nlattr *attr, *allowedip[WGALLOWEDIP_A_MAX + 1];
|
||||
int rem;
|
||||
|
||||
nla_for_each_nested(attr, attrs[WGPEER_A_ALLOWEDIPS], rem) {
|
||||
ret = nla_parse_nested(allowedip, WGALLOWEDIP_A_MAX,
|
||||
attr, allowedip_policy, NULL);
|
||||
if (ret < 0)
|
||||
goto out;
|
||||
ret = set_allowedip(peer, allowedip);
|
||||
if (ret < 0)
|
||||
goto out;
|
||||
}
|
||||
}
|
||||
|
||||
if (attrs[WGPEER_A_PERSISTENT_KEEPALIVE_INTERVAL]) {
|
||||
const u16 persistent_keepalive_interval = nla_get_u16(
|
||||
attrs[WGPEER_A_PERSISTENT_KEEPALIVE_INTERVAL]);
|
||||
const bool send_keepalive =
|
||||
!peer->persistent_keepalive_interval &&
|
||||
persistent_keepalive_interval &&
|
||||
netif_running(wg->dev);
|
||||
|
||||
peer->persistent_keepalive_interval = persistent_keepalive_interval;
|
||||
if (send_keepalive)
|
||||
wg_packet_send_keepalive(peer);
|
||||
}
|
||||
|
||||
if (netif_running(wg->dev))
|
||||
wg_packet_send_staged_packets(peer);
|
||||
|
||||
out:
|
||||
wg_peer_put(peer);
|
||||
if (attrs[WGPEER_A_PRESHARED_KEY])
|
||||
memzero_explicit(nla_data(attrs[WGPEER_A_PRESHARED_KEY]),
|
||||
nla_len(attrs[WGPEER_A_PRESHARED_KEY]));
|
||||
return ret;
|
||||
}
|
||||
|
||||
static int wg_set_device(struct sk_buff *skb, struct genl_info *info)
|
||||
{
|
||||
struct wg_device *wg = lookup_interface(info->attrs, skb);
|
||||
u32 flags = 0;
|
||||
int ret;
|
||||
|
||||
if (IS_ERR(wg)) {
|
||||
ret = PTR_ERR(wg);
|
||||
goto out_nodev;
|
||||
}
|
||||
|
||||
rtnl_lock();
|
||||
mutex_lock(&wg->device_update_lock);
|
||||
|
||||
if (info->attrs[WGDEVICE_A_FLAGS])
|
||||
flags = nla_get_u32(info->attrs[WGDEVICE_A_FLAGS]);
|
||||
ret = -EOPNOTSUPP;
|
||||
if (flags & ~__WGDEVICE_F_ALL)
|
||||
goto out;
|
||||
|
||||
if (info->attrs[WGDEVICE_A_LISTEN_PORT] || info->attrs[WGDEVICE_A_FWMARK]) {
|
||||
struct net *net;
|
||||
rcu_read_lock();
|
||||
net = rcu_dereference(wg->creating_net);
|
||||
ret = !net || !ns_capable(net->user_ns, CAP_NET_ADMIN) ? -EPERM : 0;
|
||||
rcu_read_unlock();
|
||||
if (ret)
|
||||
goto out;
|
||||
}
|
||||
|
||||
++wg->device_update_gen;
|
||||
|
||||
if (info->attrs[WGDEVICE_A_FWMARK]) {
|
||||
struct wg_peer *peer;
|
||||
|
||||
wg->fwmark = nla_get_u32(info->attrs[WGDEVICE_A_FWMARK]);
|
||||
list_for_each_entry(peer, &wg->peer_list, peer_list)
|
||||
wg_socket_clear_peer_endpoint_src(peer);
|
||||
}
|
||||
|
||||
if (info->attrs[WGDEVICE_A_LISTEN_PORT]) {
|
||||
ret = set_port(wg,
|
||||
nla_get_u16(info->attrs[WGDEVICE_A_LISTEN_PORT]));
|
||||
if (ret)
|
||||
goto out;
|
||||
}
|
||||
|
||||
if (flags & WGDEVICE_F_REPLACE_PEERS)
|
||||
wg_peer_remove_all(wg);
|
||||
|
||||
if (info->attrs[WGDEVICE_A_PRIVATE_KEY] &&
|
||||
nla_len(info->attrs[WGDEVICE_A_PRIVATE_KEY]) ==
|
||||
NOISE_PUBLIC_KEY_LEN) {
|
||||
u8 *private_key = nla_data(info->attrs[WGDEVICE_A_PRIVATE_KEY]);
|
||||
u8 public_key[NOISE_PUBLIC_KEY_LEN];
|
||||
struct wg_peer *peer, *temp;
|
||||
bool send_staged_packets;
|
||||
|
||||
if (!crypto_memneq(wg->static_identity.static_private,
|
||||
private_key, NOISE_PUBLIC_KEY_LEN))
|
||||
goto skip_set_private_key;
|
||||
|
||||
/* We remove before setting, to prevent race, which means doing
|
||||
* two 25519-genpub ops.
|
||||
*/
|
||||
if (curve25519_generate_public(public_key, private_key)) {
|
||||
peer = wg_pubkey_hashtable_lookup(wg->peer_hashtable,
|
||||
public_key);
|
||||
if (peer) {
|
||||
wg_peer_put(peer);
|
||||
wg_peer_remove(peer);
|
||||
}
|
||||
}
|
||||
|
||||
down_write(&wg->static_identity.lock);
|
||||
send_staged_packets = !wg->static_identity.has_identity && netif_running(wg->dev);
|
||||
wg_noise_set_static_identity_private_key(&wg->static_identity, private_key);
|
||||
send_staged_packets = send_staged_packets && wg->static_identity.has_identity;
|
||||
|
||||
wg_cookie_checker_precompute_device_keys(&wg->cookie_checker);
|
||||
list_for_each_entry_safe(peer, temp, &wg->peer_list, peer_list) {
|
||||
wg_noise_precompute_static_static(peer);
|
||||
wg_noise_expire_current_peer_keypairs(peer);
|
||||
if (send_staged_packets)
|
||||
wg_packet_send_staged_packets(peer);
|
||||
}
|
||||
up_write(&wg->static_identity.lock);
|
||||
}
|
||||
skip_set_private_key:
|
||||
|
||||
if (info->attrs[WGDEVICE_A_PEERS]) {
|
||||
struct nlattr *attr, *peer[WGPEER_A_MAX + 1];
|
||||
int rem;
|
||||
|
||||
nla_for_each_nested(attr, info->attrs[WGDEVICE_A_PEERS], rem) {
|
||||
ret = nla_parse_nested(peer, WGPEER_A_MAX, attr,
|
||||
peer_policy, NULL);
|
||||
if (ret < 0)
|
||||
goto out;
|
||||
ret = set_peer(wg, peer);
|
||||
if (ret < 0)
|
||||
goto out;
|
||||
}
|
||||
}
|
||||
ret = 0;
|
||||
|
||||
out:
|
||||
mutex_unlock(&wg->device_update_lock);
|
||||
rtnl_unlock();
|
||||
dev_put(wg->dev);
|
||||
out_nodev:
|
||||
if (info->attrs[WGDEVICE_A_PRIVATE_KEY])
|
||||
memzero_explicit(nla_data(info->attrs[WGDEVICE_A_PRIVATE_KEY]),
|
||||
nla_len(info->attrs[WGDEVICE_A_PRIVATE_KEY]));
|
||||
return ret;
|
||||
}
|
||||
|
||||
static const struct genl_ops genl_ops[] = {
|
||||
{
|
||||
.cmd = WG_CMD_GET_DEVICE,
|
||||
.start = wg_get_device_start,
|
||||
.dumpit = wg_get_device_dump,
|
||||
.done = wg_get_device_done,
|
||||
.flags = GENL_UNS_ADMIN_PERM
|
||||
}, {
|
||||
.cmd = WG_CMD_SET_DEVICE,
|
||||
.doit = wg_set_device,
|
||||
.flags = GENL_UNS_ADMIN_PERM
|
||||
}
|
||||
};
|
||||
|
||||
static struct genl_family genl_family __ro_after_init = {
|
||||
.ops = genl_ops,
|
||||
.n_ops = ARRAY_SIZE(genl_ops),
|
||||
.name = WG_GENL_NAME,
|
||||
.version = WG_GENL_VERSION,
|
||||
.maxattr = WGDEVICE_A_MAX,
|
||||
.module = THIS_MODULE,
|
||||
.policy = device_policy,
|
||||
.netnsok = true
|
||||
};
|
||||
|
||||
int __init wg_genetlink_init(void)
|
||||
{
|
||||
return genl_register_family(&genl_family);
|
||||
}
|
||||
|
||||
void __exit wg_genetlink_uninit(void)
|
||||
{
|
||||
genl_unregister_family(&genl_family);
|
||||
}
|
||||
12
drivers/net/wireguard/netlink.h
Normal file
12
drivers/net/wireguard/netlink.h
Normal file
|
|
@ -0,0 +1,12 @@
|
|||
/* SPDX-License-Identifier: GPL-2.0 */
|
||||
/*
|
||||
* Copyright (C) 2015-2019 Jason A. Donenfeld <Jason@zx2c4.com>. All Rights Reserved.
|
||||
*/
|
||||
|
||||
#ifndef _WG_NETLINK_H
|
||||
#define _WG_NETLINK_H
|
||||
|
||||
int wg_genetlink_init(void);
|
||||
void wg_genetlink_uninit(void);
|
||||
|
||||
#endif /* _WG_NETLINK_H */
|
||||
861
drivers/net/wireguard/noise.c
Normal file
861
drivers/net/wireguard/noise.c
Normal file
|
|
@ -0,0 +1,861 @@
|
|||
// SPDX-License-Identifier: GPL-2.0
|
||||
/*
|
||||
* Copyright (C) 2015-2019 Jason A. Donenfeld <Jason@zx2c4.com>. All Rights Reserved.
|
||||
*/
|
||||
|
||||
#include "noise.h"
|
||||
#include "device.h"
|
||||
#include "peer.h"
|
||||
#include "messages.h"
|
||||
#include "queueing.h"
|
||||
#include "peerlookup.h"
|
||||
|
||||
#include <linux/rcupdate.h>
|
||||
#include <linux/slab.h>
|
||||
#include <linux/bitmap.h>
|
||||
#include <linux/scatterlist.h>
|
||||
#include <linux/highmem.h>
|
||||
#include <crypto/algapi.h>
|
||||
|
||||
/* This implements Noise_IKpsk2:
|
||||
*
|
||||
* <- s
|
||||
* ******
|
||||
* -> e, es, s, ss, {t}
|
||||
* <- e, ee, se, psk, {}
|
||||
*/
|
||||
|
||||
static const u8 handshake_name[37] = "Noise_IKpsk2_25519_ChaChaPoly_BLAKE2s";
|
||||
static const u8 identifier_name[34] = "WireGuard v1 zx2c4 Jason@zx2c4.com";
|
||||
static u8 handshake_init_hash[NOISE_HASH_LEN] __ro_after_init;
|
||||
static u8 handshake_init_chaining_key[NOISE_HASH_LEN] __ro_after_init;
|
||||
static atomic64_t keypair_counter = ATOMIC64_INIT(0);
|
||||
|
||||
void __init wg_noise_init(void)
|
||||
{
|
||||
struct blake2s_state blake;
|
||||
|
||||
blake2s(handshake_init_chaining_key, handshake_name, NULL,
|
||||
NOISE_HASH_LEN, sizeof(handshake_name), 0);
|
||||
blake2s_init(&blake, NOISE_HASH_LEN);
|
||||
blake2s_update(&blake, handshake_init_chaining_key, NOISE_HASH_LEN);
|
||||
blake2s_update(&blake, identifier_name, sizeof(identifier_name));
|
||||
blake2s_final(&blake, handshake_init_hash);
|
||||
}
|
||||
|
||||
/* Must hold peer->handshake.static_identity->lock */
|
||||
void wg_noise_precompute_static_static(struct wg_peer *peer)
|
||||
{
|
||||
down_write(&peer->handshake.lock);
|
||||
if (!peer->handshake.static_identity->has_identity ||
|
||||
!curve25519(peer->handshake.precomputed_static_static,
|
||||
peer->handshake.static_identity->static_private,
|
||||
peer->handshake.remote_static))
|
||||
memset(peer->handshake.precomputed_static_static, 0,
|
||||
NOISE_PUBLIC_KEY_LEN);
|
||||
up_write(&peer->handshake.lock);
|
||||
}
|
||||
|
||||
void wg_noise_handshake_init(struct noise_handshake *handshake,
|
||||
struct noise_static_identity *static_identity,
|
||||
const u8 peer_public_key[NOISE_PUBLIC_KEY_LEN],
|
||||
const u8 peer_preshared_key[NOISE_SYMMETRIC_KEY_LEN],
|
||||
struct wg_peer *peer)
|
||||
{
|
||||
memset(handshake, 0, sizeof(*handshake));
|
||||
init_rwsem(&handshake->lock);
|
||||
handshake->entry.type = INDEX_HASHTABLE_HANDSHAKE;
|
||||
handshake->entry.peer = peer;
|
||||
memcpy(handshake->remote_static, peer_public_key, NOISE_PUBLIC_KEY_LEN);
|
||||
if (peer_preshared_key)
|
||||
memcpy(handshake->preshared_key, peer_preshared_key,
|
||||
NOISE_SYMMETRIC_KEY_LEN);
|
||||
handshake->static_identity = static_identity;
|
||||
handshake->state = HANDSHAKE_ZEROED;
|
||||
wg_noise_precompute_static_static(peer);
|
||||
}
|
||||
|
||||
static void handshake_zero(struct noise_handshake *handshake)
|
||||
{
|
||||
memset(&handshake->ephemeral_private, 0, NOISE_PUBLIC_KEY_LEN);
|
||||
memset(&handshake->remote_ephemeral, 0, NOISE_PUBLIC_KEY_LEN);
|
||||
memset(&handshake->hash, 0, NOISE_HASH_LEN);
|
||||
memset(&handshake->chaining_key, 0, NOISE_HASH_LEN);
|
||||
handshake->remote_index = 0;
|
||||
handshake->state = HANDSHAKE_ZEROED;
|
||||
}
|
||||
|
||||
void wg_noise_handshake_clear(struct noise_handshake *handshake)
|
||||
{
|
||||
down_write(&handshake->lock);
|
||||
wg_index_hashtable_remove(
|
||||
handshake->entry.peer->device->index_hashtable,
|
||||
&handshake->entry);
|
||||
handshake_zero(handshake);
|
||||
up_write(&handshake->lock);
|
||||
}
|
||||
|
||||
static struct noise_keypair *keypair_create(struct wg_peer *peer)
|
||||
{
|
||||
struct noise_keypair *keypair = kzalloc(sizeof(*keypair), GFP_KERNEL);
|
||||
|
||||
if (unlikely(!keypair))
|
||||
return NULL;
|
||||
spin_lock_init(&keypair->receiving_counter.lock);
|
||||
keypair->internal_id = atomic64_inc_return(&keypair_counter);
|
||||
keypair->entry.type = INDEX_HASHTABLE_KEYPAIR;
|
||||
keypair->entry.peer = peer;
|
||||
kref_init(&keypair->refcount);
|
||||
return keypair;
|
||||
}
|
||||
|
||||
static void keypair_free_rcu(struct rcu_head *rcu)
|
||||
{
|
||||
kzfree(container_of(rcu, struct noise_keypair, rcu));
|
||||
}
|
||||
|
||||
static void keypair_free_kref(struct kref *kref)
|
||||
{
|
||||
struct noise_keypair *keypair =
|
||||
container_of(kref, struct noise_keypair, refcount);
|
||||
|
||||
net_dbg_ratelimited("%s: Keypair %llu destroyed for peer %llu\n",
|
||||
keypair->entry.peer->device->dev->name,
|
||||
keypair->internal_id,
|
||||
keypair->entry.peer->internal_id);
|
||||
wg_index_hashtable_remove(keypair->entry.peer->device->index_hashtable,
|
||||
&keypair->entry);
|
||||
call_rcu(&keypair->rcu, keypair_free_rcu);
|
||||
}
|
||||
|
||||
void wg_noise_keypair_put(struct noise_keypair *keypair, bool unreference_now)
|
||||
{
|
||||
if (unlikely(!keypair))
|
||||
return;
|
||||
if (unlikely(unreference_now))
|
||||
wg_index_hashtable_remove(
|
||||
keypair->entry.peer->device->index_hashtable,
|
||||
&keypair->entry);
|
||||
kref_put(&keypair->refcount, keypair_free_kref);
|
||||
}
|
||||
|
||||
struct noise_keypair *wg_noise_keypair_get(struct noise_keypair *keypair)
|
||||
{
|
||||
RCU_LOCKDEP_WARN(!rcu_read_lock_bh_held(),
|
||||
"Taking noise keypair reference without holding the RCU BH read lock");
|
||||
if (unlikely(!keypair || !kref_get_unless_zero(&keypair->refcount)))
|
||||
return NULL;
|
||||
return keypair;
|
||||
}
|
||||
|
||||
void wg_noise_keypairs_clear(struct noise_keypairs *keypairs)
|
||||
{
|
||||
struct noise_keypair *old;
|
||||
|
||||
spin_lock_bh(&keypairs->keypair_update_lock);
|
||||
|
||||
/* We zero the next_keypair before zeroing the others, so that
|
||||
* wg_noise_received_with_keypair returns early before subsequent ones
|
||||
* are zeroed.
|
||||
*/
|
||||
old = rcu_dereference_protected(keypairs->next_keypair,
|
||||
lockdep_is_held(&keypairs->keypair_update_lock));
|
||||
RCU_INIT_POINTER(keypairs->next_keypair, NULL);
|
||||
wg_noise_keypair_put(old, true);
|
||||
|
||||
old = rcu_dereference_protected(keypairs->previous_keypair,
|
||||
lockdep_is_held(&keypairs->keypair_update_lock));
|
||||
RCU_INIT_POINTER(keypairs->previous_keypair, NULL);
|
||||
wg_noise_keypair_put(old, true);
|
||||
|
||||
old = rcu_dereference_protected(keypairs->current_keypair,
|
||||
lockdep_is_held(&keypairs->keypair_update_lock));
|
||||
RCU_INIT_POINTER(keypairs->current_keypair, NULL);
|
||||
wg_noise_keypair_put(old, true);
|
||||
|
||||
spin_unlock_bh(&keypairs->keypair_update_lock);
|
||||
}
|
||||
|
||||
void wg_noise_expire_current_peer_keypairs(struct wg_peer *peer)
|
||||
{
|
||||
struct noise_keypair *keypair;
|
||||
|
||||
wg_noise_handshake_clear(&peer->handshake);
|
||||
wg_noise_reset_last_sent_handshake(&peer->last_sent_handshake);
|
||||
|
||||
spin_lock_bh(&peer->keypairs.keypair_update_lock);
|
||||
keypair = rcu_dereference_protected(peer->keypairs.next_keypair,
|
||||
lockdep_is_held(&peer->keypairs.keypair_update_lock));
|
||||
if (keypair)
|
||||
keypair->sending.is_valid = false;
|
||||
keypair = rcu_dereference_protected(peer->keypairs.current_keypair,
|
||||
lockdep_is_held(&peer->keypairs.keypair_update_lock));
|
||||
if (keypair)
|
||||
keypair->sending.is_valid = false;
|
||||
spin_unlock_bh(&peer->keypairs.keypair_update_lock);
|
||||
}
|
||||
|
||||
static void add_new_keypair(struct noise_keypairs *keypairs,
|
||||
struct noise_keypair *new_keypair)
|
||||
{
|
||||
struct noise_keypair *previous_keypair, *next_keypair, *current_keypair;
|
||||
|
||||
spin_lock_bh(&keypairs->keypair_update_lock);
|
||||
previous_keypair = rcu_dereference_protected(keypairs->previous_keypair,
|
||||
lockdep_is_held(&keypairs->keypair_update_lock));
|
||||
next_keypair = rcu_dereference_protected(keypairs->next_keypair,
|
||||
lockdep_is_held(&keypairs->keypair_update_lock));
|
||||
current_keypair = rcu_dereference_protected(keypairs->current_keypair,
|
||||
lockdep_is_held(&keypairs->keypair_update_lock));
|
||||
if (new_keypair->i_am_the_initiator) {
|
||||
/* If we're the initiator, it means we've sent a handshake, and
|
||||
* received a confirmation response, which means this new
|
||||
* keypair can now be used.
|
||||
*/
|
||||
if (next_keypair) {
|
||||
/* If there already was a next keypair pending, we
|
||||
* demote it to be the previous keypair, and free the
|
||||
* existing current. Note that this means KCI can result
|
||||
* in this transition. It would perhaps be more sound to
|
||||
* always just get rid of the unused next keypair
|
||||
* instead of putting it in the previous slot, but this
|
||||
* might be a bit less robust. Something to think about
|
||||
* for the future.
|
||||
*/
|
||||
RCU_INIT_POINTER(keypairs->next_keypair, NULL);
|
||||
rcu_assign_pointer(keypairs->previous_keypair,
|
||||
next_keypair);
|
||||
wg_noise_keypair_put(current_keypair, true);
|
||||
} else /* If there wasn't an existing next keypair, we replace
|
||||
* the previous with the current one.
|
||||
*/
|
||||
rcu_assign_pointer(keypairs->previous_keypair,
|
||||
current_keypair);
|
||||
/* At this point we can get rid of the old previous keypair, and
|
||||
* set up the new keypair.
|
||||
*/
|
||||
wg_noise_keypair_put(previous_keypair, true);
|
||||
rcu_assign_pointer(keypairs->current_keypair, new_keypair);
|
||||
} else {
|
||||
/* If we're the responder, it means we can't use the new keypair
|
||||
* until we receive confirmation via the first data packet, so
|
||||
* we get rid of the existing previous one, the possibly
|
||||
* existing next one, and slide in the new next one.
|
||||
*/
|
||||
rcu_assign_pointer(keypairs->next_keypair, new_keypair);
|
||||
wg_noise_keypair_put(next_keypair, true);
|
||||
RCU_INIT_POINTER(keypairs->previous_keypair, NULL);
|
||||
wg_noise_keypair_put(previous_keypair, true);
|
||||
}
|
||||
spin_unlock_bh(&keypairs->keypair_update_lock);
|
||||
}
|
||||
|
||||
bool wg_noise_received_with_keypair(struct noise_keypairs *keypairs,
|
||||
struct noise_keypair *received_keypair)
|
||||
{
|
||||
struct noise_keypair *old_keypair;
|
||||
bool key_is_new;
|
||||
|
||||
/* We first check without taking the spinlock. */
|
||||
key_is_new = received_keypair ==
|
||||
rcu_access_pointer(keypairs->next_keypair);
|
||||
if (likely(!key_is_new))
|
||||
return false;
|
||||
|
||||
spin_lock_bh(&keypairs->keypair_update_lock);
|
||||
/* After locking, we double check that things didn't change from
|
||||
* beneath us.
|
||||
*/
|
||||
if (unlikely(received_keypair !=
|
||||
rcu_dereference_protected(keypairs->next_keypair,
|
||||
lockdep_is_held(&keypairs->keypair_update_lock)))) {
|
||||
spin_unlock_bh(&keypairs->keypair_update_lock);
|
||||
return false;
|
||||
}
|
||||
|
||||
/* When we've finally received the confirmation, we slide the next
|
||||
* into the current, the current into the previous, and get rid of
|
||||
* the old previous.
|
||||
*/
|
||||
old_keypair = rcu_dereference_protected(keypairs->previous_keypair,
|
||||
lockdep_is_held(&keypairs->keypair_update_lock));
|
||||
rcu_assign_pointer(keypairs->previous_keypair,
|
||||
rcu_dereference_protected(keypairs->current_keypair,
|
||||
lockdep_is_held(&keypairs->keypair_update_lock)));
|
||||
wg_noise_keypair_put(old_keypair, true);
|
||||
rcu_assign_pointer(keypairs->current_keypair, received_keypair);
|
||||
RCU_INIT_POINTER(keypairs->next_keypair, NULL);
|
||||
|
||||
spin_unlock_bh(&keypairs->keypair_update_lock);
|
||||
return true;
|
||||
}
|
||||
|
||||
/* Must hold static_identity->lock */
|
||||
void wg_noise_set_static_identity_private_key(
|
||||
struct noise_static_identity *static_identity,
|
||||
const u8 private_key[NOISE_PUBLIC_KEY_LEN])
|
||||
{
|
||||
memcpy(static_identity->static_private, private_key,
|
||||
NOISE_PUBLIC_KEY_LEN);
|
||||
curve25519_clamp_secret(static_identity->static_private);
|
||||
static_identity->has_identity = curve25519_generate_public(
|
||||
static_identity->static_public, private_key);
|
||||
}
|
||||
|
||||
static void hmac(u8 *out, const u8 *in, const u8 *key, const size_t inlen, const size_t keylen)
|
||||
{
|
||||
struct blake2s_state state;
|
||||
u8 x_key[BLAKE2S_BLOCK_SIZE] __aligned(__alignof__(u32)) = { 0 };
|
||||
u8 i_hash[BLAKE2S_HASH_SIZE] __aligned(__alignof__(u32));
|
||||
int i;
|
||||
|
||||
if (keylen > BLAKE2S_BLOCK_SIZE) {
|
||||
blake2s_init(&state, BLAKE2S_HASH_SIZE);
|
||||
blake2s_update(&state, key, keylen);
|
||||
blake2s_final(&state, x_key);
|
||||
} else
|
||||
memcpy(x_key, key, keylen);
|
||||
|
||||
for (i = 0; i < BLAKE2S_BLOCK_SIZE; ++i)
|
||||
x_key[i] ^= 0x36;
|
||||
|
||||
blake2s_init(&state, BLAKE2S_HASH_SIZE);
|
||||
blake2s_update(&state, x_key, BLAKE2S_BLOCK_SIZE);
|
||||
blake2s_update(&state, in, inlen);
|
||||
blake2s_final(&state, i_hash);
|
||||
|
||||
for (i = 0; i < BLAKE2S_BLOCK_SIZE; ++i)
|
||||
x_key[i] ^= 0x5c ^ 0x36;
|
||||
|
||||
blake2s_init(&state, BLAKE2S_HASH_SIZE);
|
||||
blake2s_update(&state, x_key, BLAKE2S_BLOCK_SIZE);
|
||||
blake2s_update(&state, i_hash, BLAKE2S_HASH_SIZE);
|
||||
blake2s_final(&state, i_hash);
|
||||
|
||||
memcpy(out, i_hash, BLAKE2S_HASH_SIZE);
|
||||
memzero_explicit(x_key, BLAKE2S_BLOCK_SIZE);
|
||||
memzero_explicit(i_hash, BLAKE2S_HASH_SIZE);
|
||||
}
|
||||
|
||||
/* This is Hugo Krawczyk's HKDF:
|
||||
* - https://eprint.iacr.org/2010/264.pdf
|
||||
* - https://tools.ietf.org/html/rfc5869
|
||||
*/
|
||||
static void kdf(u8 *first_dst, u8 *second_dst, u8 *third_dst, const u8 *data,
|
||||
size_t first_len, size_t second_len, size_t third_len,
|
||||
size_t data_len, const u8 chaining_key[NOISE_HASH_LEN])
|
||||
{
|
||||
u8 output[BLAKE2S_HASH_SIZE + 1];
|
||||
u8 secret[BLAKE2S_HASH_SIZE];
|
||||
|
||||
WARN_ON(IS_ENABLED(DEBUG) &&
|
||||
(first_len > BLAKE2S_HASH_SIZE ||
|
||||
second_len > BLAKE2S_HASH_SIZE ||
|
||||
third_len > BLAKE2S_HASH_SIZE ||
|
||||
((second_len || second_dst || third_len || third_dst) &&
|
||||
(!first_len || !first_dst)) ||
|
||||
((third_len || third_dst) && (!second_len || !second_dst))));
|
||||
|
||||
/* Extract entropy from data into secret */
|
||||
hmac(secret, data, chaining_key, data_len, NOISE_HASH_LEN);
|
||||
|
||||
if (!first_dst || !first_len)
|
||||
goto out;
|
||||
|
||||
/* Expand first key: key = secret, data = 0x1 */
|
||||
output[0] = 1;
|
||||
hmac(output, output, secret, 1, BLAKE2S_HASH_SIZE);
|
||||
memcpy(first_dst, output, first_len);
|
||||
|
||||
if (!second_dst || !second_len)
|
||||
goto out;
|
||||
|
||||
/* Expand second key: key = secret, data = first-key || 0x2 */
|
||||
output[BLAKE2S_HASH_SIZE] = 2;
|
||||
hmac(output, output, secret, BLAKE2S_HASH_SIZE + 1, BLAKE2S_HASH_SIZE);
|
||||
memcpy(second_dst, output, second_len);
|
||||
|
||||
if (!third_dst || !third_len)
|
||||
goto out;
|
||||
|
||||
/* Expand third key: key = secret, data = second-key || 0x3 */
|
||||
output[BLAKE2S_HASH_SIZE] = 3;
|
||||
hmac(output, output, secret, BLAKE2S_HASH_SIZE + 1, BLAKE2S_HASH_SIZE);
|
||||
memcpy(third_dst, output, third_len);
|
||||
|
||||
out:
|
||||
/* Clear sensitive data from stack */
|
||||
memzero_explicit(secret, BLAKE2S_HASH_SIZE);
|
||||
memzero_explicit(output, BLAKE2S_HASH_SIZE + 1);
|
||||
}
|
||||
|
||||
static void derive_keys(struct noise_symmetric_key *first_dst,
|
||||
struct noise_symmetric_key *second_dst,
|
||||
const u8 chaining_key[NOISE_HASH_LEN])
|
||||
{
|
||||
u64 birthdate = ktime_get_coarse_boottime_ns();
|
||||
kdf(first_dst->key, second_dst->key, NULL, NULL,
|
||||
NOISE_SYMMETRIC_KEY_LEN, NOISE_SYMMETRIC_KEY_LEN, 0, 0,
|
||||
chaining_key);
|
||||
first_dst->birthdate = second_dst->birthdate = birthdate;
|
||||
first_dst->is_valid = second_dst->is_valid = true;
|
||||
}
|
||||
|
||||
static bool __must_check mix_dh(u8 chaining_key[NOISE_HASH_LEN],
|
||||
u8 key[NOISE_SYMMETRIC_KEY_LEN],
|
||||
const u8 private[NOISE_PUBLIC_KEY_LEN],
|
||||
const u8 public[NOISE_PUBLIC_KEY_LEN])
|
||||
{
|
||||
u8 dh_calculation[NOISE_PUBLIC_KEY_LEN];
|
||||
|
||||
if (unlikely(!curve25519(dh_calculation, private, public)))
|
||||
return false;
|
||||
kdf(chaining_key, key, NULL, dh_calculation, NOISE_HASH_LEN,
|
||||
NOISE_SYMMETRIC_KEY_LEN, 0, NOISE_PUBLIC_KEY_LEN, chaining_key);
|
||||
memzero_explicit(dh_calculation, NOISE_PUBLIC_KEY_LEN);
|
||||
return true;
|
||||
}
|
||||
|
||||
static bool __must_check mix_precomputed_dh(u8 chaining_key[NOISE_HASH_LEN],
|
||||
u8 key[NOISE_SYMMETRIC_KEY_LEN],
|
||||
const u8 precomputed[NOISE_PUBLIC_KEY_LEN])
|
||||
{
|
||||
static u8 zero_point[NOISE_PUBLIC_KEY_LEN];
|
||||
if (unlikely(!crypto_memneq(precomputed, zero_point, NOISE_PUBLIC_KEY_LEN)))
|
||||
return false;
|
||||
kdf(chaining_key, key, NULL, precomputed, NOISE_HASH_LEN,
|
||||
NOISE_SYMMETRIC_KEY_LEN, 0, NOISE_PUBLIC_KEY_LEN,
|
||||
chaining_key);
|
||||
return true;
|
||||
}
|
||||
|
||||
static void mix_hash(u8 hash[NOISE_HASH_LEN], const u8 *src, size_t src_len)
|
||||
{
|
||||
struct blake2s_state blake;
|
||||
|
||||
blake2s_init(&blake, NOISE_HASH_LEN);
|
||||
blake2s_update(&blake, hash, NOISE_HASH_LEN);
|
||||
blake2s_update(&blake, src, src_len);
|
||||
blake2s_final(&blake, hash);
|
||||
}
|
||||
|
||||
static void mix_psk(u8 chaining_key[NOISE_HASH_LEN], u8 hash[NOISE_HASH_LEN],
|
||||
u8 key[NOISE_SYMMETRIC_KEY_LEN],
|
||||
const u8 psk[NOISE_SYMMETRIC_KEY_LEN])
|
||||
{
|
||||
u8 temp_hash[NOISE_HASH_LEN];
|
||||
|
||||
kdf(chaining_key, temp_hash, key, psk, NOISE_HASH_LEN, NOISE_HASH_LEN,
|
||||
NOISE_SYMMETRIC_KEY_LEN, NOISE_SYMMETRIC_KEY_LEN, chaining_key);
|
||||
mix_hash(hash, temp_hash, NOISE_HASH_LEN);
|
||||
memzero_explicit(temp_hash, NOISE_HASH_LEN);
|
||||
}
|
||||
|
||||
static void handshake_init(u8 chaining_key[NOISE_HASH_LEN],
|
||||
u8 hash[NOISE_HASH_LEN],
|
||||
const u8 remote_static[NOISE_PUBLIC_KEY_LEN])
|
||||
{
|
||||
memcpy(hash, handshake_init_hash, NOISE_HASH_LEN);
|
||||
memcpy(chaining_key, handshake_init_chaining_key, NOISE_HASH_LEN);
|
||||
mix_hash(hash, remote_static, NOISE_PUBLIC_KEY_LEN);
|
||||
}
|
||||
|
||||
static void message_encrypt(u8 *dst_ciphertext, const u8 *src_plaintext,
|
||||
size_t src_len, u8 key[NOISE_SYMMETRIC_KEY_LEN],
|
||||
u8 hash[NOISE_HASH_LEN])
|
||||
{
|
||||
chacha20poly1305_encrypt(dst_ciphertext, src_plaintext, src_len, hash,
|
||||
NOISE_HASH_LEN,
|
||||
0 /* Always zero for Noise_IK */, key);
|
||||
mix_hash(hash, dst_ciphertext, noise_encrypted_len(src_len));
|
||||
}
|
||||
|
||||
static bool message_decrypt(u8 *dst_plaintext, const u8 *src_ciphertext,
|
||||
size_t src_len, u8 key[NOISE_SYMMETRIC_KEY_LEN],
|
||||
u8 hash[NOISE_HASH_LEN])
|
||||
{
|
||||
if (!chacha20poly1305_decrypt(dst_plaintext, src_ciphertext, src_len,
|
||||
hash, NOISE_HASH_LEN,
|
||||
0 /* Always zero for Noise_IK */, key))
|
||||
return false;
|
||||
mix_hash(hash, src_ciphertext, src_len);
|
||||
return true;
|
||||
}
|
||||
|
||||
static void message_ephemeral(u8 ephemeral_dst[NOISE_PUBLIC_KEY_LEN],
|
||||
const u8 ephemeral_src[NOISE_PUBLIC_KEY_LEN],
|
||||
u8 chaining_key[NOISE_HASH_LEN],
|
||||
u8 hash[NOISE_HASH_LEN])
|
||||
{
|
||||
if (ephemeral_dst != ephemeral_src)
|
||||
memcpy(ephemeral_dst, ephemeral_src, NOISE_PUBLIC_KEY_LEN);
|
||||
mix_hash(hash, ephemeral_src, NOISE_PUBLIC_KEY_LEN);
|
||||
kdf(chaining_key, NULL, NULL, ephemeral_src, NOISE_HASH_LEN, 0, 0,
|
||||
NOISE_PUBLIC_KEY_LEN, chaining_key);
|
||||
}
|
||||
|
||||
static void tai64n_now(u8 output[NOISE_TIMESTAMP_LEN])
|
||||
{
|
||||
struct timespec64 now;
|
||||
|
||||
ktime_get_real_ts64(&now);
|
||||
|
||||
/* In order to prevent some sort of infoleak from precise timers, we
|
||||
* round down the nanoseconds part to the closest rounded-down power of
|
||||
* two to the maximum initiations per second allowed anyway by the
|
||||
* implementation.
|
||||
*/
|
||||
now.tv_nsec = ALIGN_DOWN(now.tv_nsec,
|
||||
rounddown_pow_of_two(NSEC_PER_SEC / INITIATIONS_PER_SECOND));
|
||||
|
||||
/* https://cr.yp.to/libtai/tai64.html */
|
||||
*(__be64 *)output = cpu_to_be64(0x400000000000000aULL + now.tv_sec);
|
||||
*(__be32 *)(output + sizeof(__be64)) = cpu_to_be32(now.tv_nsec);
|
||||
}
|
||||
|
||||
bool
|
||||
wg_noise_handshake_create_initiation(struct message_handshake_initiation *dst,
|
||||
struct noise_handshake *handshake)
|
||||
{
|
||||
u8 timestamp[NOISE_TIMESTAMP_LEN];
|
||||
u8 key[NOISE_SYMMETRIC_KEY_LEN];
|
||||
bool ret = false;
|
||||
|
||||
/* We need to wait for crng _before_ taking any locks, since
|
||||
* curve25519_generate_secret uses get_random_bytes_wait.
|
||||
*/
|
||||
wait_for_random_bytes();
|
||||
|
||||
down_read(&handshake->static_identity->lock);
|
||||
down_write(&handshake->lock);
|
||||
|
||||
if (unlikely(!handshake->static_identity->has_identity))
|
||||
goto out;
|
||||
|
||||
dst->header.type = cpu_to_le32(MESSAGE_HANDSHAKE_INITIATION);
|
||||
|
||||
handshake_init(handshake->chaining_key, handshake->hash,
|
||||
handshake->remote_static);
|
||||
|
||||
/* e */
|
||||
curve25519_generate_secret(handshake->ephemeral_private);
|
||||
if (!curve25519_generate_public(dst->unencrypted_ephemeral,
|
||||
handshake->ephemeral_private))
|
||||
goto out;
|
||||
message_ephemeral(dst->unencrypted_ephemeral,
|
||||
dst->unencrypted_ephemeral, handshake->chaining_key,
|
||||
handshake->hash);
|
||||
|
||||
/* es */
|
||||
if (!mix_dh(handshake->chaining_key, key, handshake->ephemeral_private,
|
||||
handshake->remote_static))
|
||||
goto out;
|
||||
|
||||
/* s */
|
||||
message_encrypt(dst->encrypted_static,
|
||||
handshake->static_identity->static_public,
|
||||
NOISE_PUBLIC_KEY_LEN, key, handshake->hash);
|
||||
|
||||
/* ss */
|
||||
if (!mix_precomputed_dh(handshake->chaining_key, key,
|
||||
handshake->precomputed_static_static))
|
||||
goto out;
|
||||
|
||||
/* {t} */
|
||||
tai64n_now(timestamp);
|
||||
message_encrypt(dst->encrypted_timestamp, timestamp,
|
||||
NOISE_TIMESTAMP_LEN, key, handshake->hash);
|
||||
|
||||
dst->sender_index = wg_index_hashtable_insert(
|
||||
handshake->entry.peer->device->index_hashtable,
|
||||
&handshake->entry);
|
||||
|
||||
handshake->state = HANDSHAKE_CREATED_INITIATION;
|
||||
ret = true;
|
||||
|
||||
out:
|
||||
up_write(&handshake->lock);
|
||||
up_read(&handshake->static_identity->lock);
|
||||
memzero_explicit(key, NOISE_SYMMETRIC_KEY_LEN);
|
||||
return ret;
|
||||
}
|
||||
|
||||
struct wg_peer *
|
||||
wg_noise_handshake_consume_initiation(struct message_handshake_initiation *src,
|
||||
struct wg_device *wg)
|
||||
{
|
||||
struct wg_peer *peer = NULL, *ret_peer = NULL;
|
||||
struct noise_handshake *handshake;
|
||||
bool replay_attack, flood_attack;
|
||||
u8 key[NOISE_SYMMETRIC_KEY_LEN];
|
||||
u8 chaining_key[NOISE_HASH_LEN];
|
||||
u8 hash[NOISE_HASH_LEN];
|
||||
u8 s[NOISE_PUBLIC_KEY_LEN];
|
||||
u8 e[NOISE_PUBLIC_KEY_LEN];
|
||||
u8 t[NOISE_TIMESTAMP_LEN];
|
||||
u64 initiation_consumption;
|
||||
|
||||
down_read(&wg->static_identity.lock);
|
||||
if (unlikely(!wg->static_identity.has_identity))
|
||||
goto out;
|
||||
|
||||
handshake_init(chaining_key, hash, wg->static_identity.static_public);
|
||||
|
||||
/* e */
|
||||
message_ephemeral(e, src->unencrypted_ephemeral, chaining_key, hash);
|
||||
|
||||
/* es */
|
||||
if (!mix_dh(chaining_key, key, wg->static_identity.static_private, e))
|
||||
goto out;
|
||||
|
||||
/* s */
|
||||
if (!message_decrypt(s, src->encrypted_static,
|
||||
sizeof(src->encrypted_static), key, hash))
|
||||
goto out;
|
||||
|
||||
/* Lookup which peer we're actually talking to */
|
||||
peer = wg_pubkey_hashtable_lookup(wg->peer_hashtable, s);
|
||||
if (!peer)
|
||||
goto out;
|
||||
handshake = &peer->handshake;
|
||||
|
||||
/* ss */
|
||||
if (!mix_precomputed_dh(chaining_key, key,
|
||||
handshake->precomputed_static_static))
|
||||
goto out;
|
||||
|
||||
/* {t} */
|
||||
if (!message_decrypt(t, src->encrypted_timestamp,
|
||||
sizeof(src->encrypted_timestamp), key, hash))
|
||||
goto out;
|
||||
|
||||
down_read(&handshake->lock);
|
||||
replay_attack = memcmp(t, handshake->latest_timestamp,
|
||||
NOISE_TIMESTAMP_LEN) <= 0;
|
||||
flood_attack = (s64)handshake->last_initiation_consumption +
|
||||
NSEC_PER_SEC / INITIATIONS_PER_SECOND >
|
||||
(s64)ktime_get_coarse_boottime_ns();
|
||||
up_read(&handshake->lock);
|
||||
if (replay_attack || flood_attack)
|
||||
goto out;
|
||||
|
||||
/* Success! Copy everything to peer */
|
||||
down_write(&handshake->lock);
|
||||
memcpy(handshake->remote_ephemeral, e, NOISE_PUBLIC_KEY_LEN);
|
||||
if (memcmp(t, handshake->latest_timestamp, NOISE_TIMESTAMP_LEN) > 0)
|
||||
memcpy(handshake->latest_timestamp, t, NOISE_TIMESTAMP_LEN);
|
||||
memcpy(handshake->hash, hash, NOISE_HASH_LEN);
|
||||
memcpy(handshake->chaining_key, chaining_key, NOISE_HASH_LEN);
|
||||
handshake->remote_index = src->sender_index;
|
||||
initiation_consumption = ktime_get_coarse_boottime_ns();
|
||||
if ((s64)(handshake->last_initiation_consumption - initiation_consumption) < 0)
|
||||
handshake->last_initiation_consumption = initiation_consumption;
|
||||
handshake->state = HANDSHAKE_CONSUMED_INITIATION;
|
||||
up_write(&handshake->lock);
|
||||
ret_peer = peer;
|
||||
|
||||
out:
|
||||
memzero_explicit(key, NOISE_SYMMETRIC_KEY_LEN);
|
||||
memzero_explicit(hash, NOISE_HASH_LEN);
|
||||
memzero_explicit(chaining_key, NOISE_HASH_LEN);
|
||||
up_read(&wg->static_identity.lock);
|
||||
if (!ret_peer)
|
||||
wg_peer_put(peer);
|
||||
return ret_peer;
|
||||
}
|
||||
|
||||
bool wg_noise_handshake_create_response(struct message_handshake_response *dst,
|
||||
struct noise_handshake *handshake)
|
||||
{
|
||||
u8 key[NOISE_SYMMETRIC_KEY_LEN];
|
||||
bool ret = false;
|
||||
|
||||
/* We need to wait for crng _before_ taking any locks, since
|
||||
* curve25519_generate_secret uses get_random_bytes_wait.
|
||||
*/
|
||||
wait_for_random_bytes();
|
||||
|
||||
down_read(&handshake->static_identity->lock);
|
||||
down_write(&handshake->lock);
|
||||
|
||||
if (handshake->state != HANDSHAKE_CONSUMED_INITIATION)
|
||||
goto out;
|
||||
|
||||
dst->header.type = cpu_to_le32(MESSAGE_HANDSHAKE_RESPONSE);
|
||||
dst->receiver_index = handshake->remote_index;
|
||||
|
||||
/* e */
|
||||
curve25519_generate_secret(handshake->ephemeral_private);
|
||||
if (!curve25519_generate_public(dst->unencrypted_ephemeral,
|
||||
handshake->ephemeral_private))
|
||||
goto out;
|
||||
message_ephemeral(dst->unencrypted_ephemeral,
|
||||
dst->unencrypted_ephemeral, handshake->chaining_key,
|
||||
handshake->hash);
|
||||
|
||||
/* ee */
|
||||
if (!mix_dh(handshake->chaining_key, NULL, handshake->ephemeral_private,
|
||||
handshake->remote_ephemeral))
|
||||
goto out;
|
||||
|
||||
/* se */
|
||||
if (!mix_dh(handshake->chaining_key, NULL, handshake->ephemeral_private,
|
||||
handshake->remote_static))
|
||||
goto out;
|
||||
|
||||
/* psk */
|
||||
mix_psk(handshake->chaining_key, handshake->hash, key,
|
||||
handshake->preshared_key);
|
||||
|
||||
/* {} */
|
||||
message_encrypt(dst->encrypted_nothing, NULL, 0, key, handshake->hash);
|
||||
|
||||
dst->sender_index = wg_index_hashtable_insert(
|
||||
handshake->entry.peer->device->index_hashtable,
|
||||
&handshake->entry);
|
||||
|
||||
handshake->state = HANDSHAKE_CREATED_RESPONSE;
|
||||
ret = true;
|
||||
|
||||
out:
|
||||
up_write(&handshake->lock);
|
||||
up_read(&handshake->static_identity->lock);
|
||||
memzero_explicit(key, NOISE_SYMMETRIC_KEY_LEN);
|
||||
return ret;
|
||||
}
|
||||
|
||||
struct wg_peer *
|
||||
wg_noise_handshake_consume_response(struct message_handshake_response *src,
|
||||
struct wg_device *wg)
|
||||
{
|
||||
enum noise_handshake_state state = HANDSHAKE_ZEROED;
|
||||
struct wg_peer *peer = NULL, *ret_peer = NULL;
|
||||
struct noise_handshake *handshake;
|
||||
u8 key[NOISE_SYMMETRIC_KEY_LEN];
|
||||
u8 hash[NOISE_HASH_LEN];
|
||||
u8 chaining_key[NOISE_HASH_LEN];
|
||||
u8 e[NOISE_PUBLIC_KEY_LEN];
|
||||
u8 ephemeral_private[NOISE_PUBLIC_KEY_LEN];
|
||||
u8 static_private[NOISE_PUBLIC_KEY_LEN];
|
||||
u8 preshared_key[NOISE_SYMMETRIC_KEY_LEN];
|
||||
|
||||
down_read(&wg->static_identity.lock);
|
||||
|
||||
if (unlikely(!wg->static_identity.has_identity))
|
||||
goto out;
|
||||
|
||||
handshake = (struct noise_handshake *)wg_index_hashtable_lookup(
|
||||
wg->index_hashtable, INDEX_HASHTABLE_HANDSHAKE,
|
||||
src->receiver_index, &peer);
|
||||
if (unlikely(!handshake))
|
||||
goto out;
|
||||
|
||||
down_read(&handshake->lock);
|
||||
state = handshake->state;
|
||||
memcpy(hash, handshake->hash, NOISE_HASH_LEN);
|
||||
memcpy(chaining_key, handshake->chaining_key, NOISE_HASH_LEN);
|
||||
memcpy(ephemeral_private, handshake->ephemeral_private,
|
||||
NOISE_PUBLIC_KEY_LEN);
|
||||
memcpy(preshared_key, handshake->preshared_key,
|
||||
NOISE_SYMMETRIC_KEY_LEN);
|
||||
up_read(&handshake->lock);
|
||||
|
||||
if (state != HANDSHAKE_CREATED_INITIATION)
|
||||
goto fail;
|
||||
|
||||
/* e */
|
||||
message_ephemeral(e, src->unencrypted_ephemeral, chaining_key, hash);
|
||||
|
||||
/* ee */
|
||||
if (!mix_dh(chaining_key, NULL, ephemeral_private, e))
|
||||
goto fail;
|
||||
|
||||
/* se */
|
||||
if (!mix_dh(chaining_key, NULL, wg->static_identity.static_private, e))
|
||||
goto fail;
|
||||
|
||||
/* psk */
|
||||
mix_psk(chaining_key, hash, key, preshared_key);
|
||||
|
||||
/* {} */
|
||||
if (!message_decrypt(NULL, src->encrypted_nothing,
|
||||
sizeof(src->encrypted_nothing), key, hash))
|
||||
goto fail;
|
||||
|
||||
/* Success! Copy everything to peer */
|
||||
down_write(&handshake->lock);
|
||||
/* It's important to check that the state is still the same, while we
|
||||
* have an exclusive lock.
|
||||
*/
|
||||
if (handshake->state != state) {
|
||||
up_write(&handshake->lock);
|
||||
goto fail;
|
||||
}
|
||||
memcpy(handshake->remote_ephemeral, e, NOISE_PUBLIC_KEY_LEN);
|
||||
memcpy(handshake->hash, hash, NOISE_HASH_LEN);
|
||||
memcpy(handshake->chaining_key, chaining_key, NOISE_HASH_LEN);
|
||||
handshake->remote_index = src->sender_index;
|
||||
handshake->state = HANDSHAKE_CONSUMED_RESPONSE;
|
||||
up_write(&handshake->lock);
|
||||
ret_peer = peer;
|
||||
goto out;
|
||||
|
||||
fail:
|
||||
wg_peer_put(peer);
|
||||
out:
|
||||
memzero_explicit(key, NOISE_SYMMETRIC_KEY_LEN);
|
||||
memzero_explicit(hash, NOISE_HASH_LEN);
|
||||
memzero_explicit(chaining_key, NOISE_HASH_LEN);
|
||||
memzero_explicit(ephemeral_private, NOISE_PUBLIC_KEY_LEN);
|
||||
memzero_explicit(static_private, NOISE_PUBLIC_KEY_LEN);
|
||||
memzero_explicit(preshared_key, NOISE_SYMMETRIC_KEY_LEN);
|
||||
up_read(&wg->static_identity.lock);
|
||||
return ret_peer;
|
||||
}
|
||||
|
||||
bool wg_noise_handshake_begin_session(struct noise_handshake *handshake,
|
||||
struct noise_keypairs *keypairs)
|
||||
{
|
||||
struct noise_keypair *new_keypair;
|
||||
bool ret = false;
|
||||
|
||||
down_write(&handshake->lock);
|
||||
if (handshake->state != HANDSHAKE_CREATED_RESPONSE &&
|
||||
handshake->state != HANDSHAKE_CONSUMED_RESPONSE)
|
||||
goto out;
|
||||
|
||||
new_keypair = keypair_create(handshake->entry.peer);
|
||||
if (!new_keypair)
|
||||
goto out;
|
||||
new_keypair->i_am_the_initiator = handshake->state ==
|
||||
HANDSHAKE_CONSUMED_RESPONSE;
|
||||
new_keypair->remote_index = handshake->remote_index;
|
||||
|
||||
if (new_keypair->i_am_the_initiator)
|
||||
derive_keys(&new_keypair->sending, &new_keypair->receiving,
|
||||
handshake->chaining_key);
|
||||
else
|
||||
derive_keys(&new_keypair->receiving, &new_keypair->sending,
|
||||
handshake->chaining_key);
|
||||
|
||||
handshake_zero(handshake);
|
||||
rcu_read_lock_bh();
|
||||
if (likely(!READ_ONCE(container_of(handshake, struct wg_peer,
|
||||
handshake)->is_dead))) {
|
||||
add_new_keypair(keypairs, new_keypair);
|
||||
net_dbg_ratelimited("%s: Keypair %llu created for peer %llu\n",
|
||||
handshake->entry.peer->device->dev->name,
|
||||
new_keypair->internal_id,
|
||||
handshake->entry.peer->internal_id);
|
||||
ret = wg_index_hashtable_replace(
|
||||
handshake->entry.peer->device->index_hashtable,
|
||||
&handshake->entry, &new_keypair->entry);
|
||||
} else {
|
||||
kzfree(new_keypair);
|
||||
}
|
||||
rcu_read_unlock_bh();
|
||||
|
||||
out:
|
||||
up_write(&handshake->lock);
|
||||
return ret;
|
||||
}
|
||||
135
drivers/net/wireguard/noise.h
Normal file
135
drivers/net/wireguard/noise.h
Normal file
|
|
@ -0,0 +1,135 @@
|
|||
/* SPDX-License-Identifier: GPL-2.0 */
|
||||
/*
|
||||
* Copyright (C) 2015-2019 Jason A. Donenfeld <Jason@zx2c4.com>. All Rights Reserved.
|
||||
*/
|
||||
#ifndef _WG_NOISE_H
|
||||
#define _WG_NOISE_H
|
||||
|
||||
#include "messages.h"
|
||||
#include "peerlookup.h"
|
||||
|
||||
#include <linux/types.h>
|
||||
#include <linux/spinlock.h>
|
||||
#include <linux/atomic.h>
|
||||
#include <linux/rwsem.h>
|
||||
#include <linux/mutex.h>
|
||||
#include <linux/kref.h>
|
||||
|
||||
struct noise_replay_counter {
|
||||
u64 counter;
|
||||
spinlock_t lock;
|
||||
unsigned long backtrack[COUNTER_BITS_TOTAL / BITS_PER_LONG];
|
||||
};
|
||||
|
||||
struct noise_symmetric_key {
|
||||
u8 key[NOISE_SYMMETRIC_KEY_LEN];
|
||||
u64 birthdate;
|
||||
bool is_valid;
|
||||
};
|
||||
|
||||
struct noise_keypair {
|
||||
struct index_hashtable_entry entry;
|
||||
struct noise_symmetric_key sending;
|
||||
atomic64_t sending_counter;
|
||||
struct noise_symmetric_key receiving;
|
||||
struct noise_replay_counter receiving_counter;
|
||||
__le32 remote_index;
|
||||
bool i_am_the_initiator;
|
||||
struct kref refcount;
|
||||
struct rcu_head rcu;
|
||||
u64 internal_id;
|
||||
};
|
||||
|
||||
struct noise_keypairs {
|
||||
struct noise_keypair __rcu *current_keypair;
|
||||
struct noise_keypair __rcu *previous_keypair;
|
||||
struct noise_keypair __rcu *next_keypair;
|
||||
spinlock_t keypair_update_lock;
|
||||
};
|
||||
|
||||
struct noise_static_identity {
|
||||
u8 static_public[NOISE_PUBLIC_KEY_LEN];
|
||||
u8 static_private[NOISE_PUBLIC_KEY_LEN];
|
||||
struct rw_semaphore lock;
|
||||
bool has_identity;
|
||||
};
|
||||
|
||||
enum noise_handshake_state {
|
||||
HANDSHAKE_ZEROED,
|
||||
HANDSHAKE_CREATED_INITIATION,
|
||||
HANDSHAKE_CONSUMED_INITIATION,
|
||||
HANDSHAKE_CREATED_RESPONSE,
|
||||
HANDSHAKE_CONSUMED_RESPONSE
|
||||
};
|
||||
|
||||
struct noise_handshake {
|
||||
struct index_hashtable_entry entry;
|
||||
|
||||
enum noise_handshake_state state;
|
||||
u64 last_initiation_consumption;
|
||||
|
||||
struct noise_static_identity *static_identity;
|
||||
|
||||
u8 ephemeral_private[NOISE_PUBLIC_KEY_LEN];
|
||||
u8 remote_static[NOISE_PUBLIC_KEY_LEN];
|
||||
u8 remote_ephemeral[NOISE_PUBLIC_KEY_LEN];
|
||||
u8 precomputed_static_static[NOISE_PUBLIC_KEY_LEN];
|
||||
|
||||
u8 preshared_key[NOISE_SYMMETRIC_KEY_LEN];
|
||||
|
||||
u8 hash[NOISE_HASH_LEN];
|
||||
u8 chaining_key[NOISE_HASH_LEN];
|
||||
|
||||
u8 latest_timestamp[NOISE_TIMESTAMP_LEN];
|
||||
__le32 remote_index;
|
||||
|
||||
/* Protects all members except the immutable (after noise_handshake_
|
||||
* init): remote_static, precomputed_static_static, static_identity.
|
||||
*/
|
||||
struct rw_semaphore lock;
|
||||
};
|
||||
|
||||
struct wg_device;
|
||||
|
||||
void wg_noise_init(void);
|
||||
void wg_noise_handshake_init(struct noise_handshake *handshake,
|
||||
struct noise_static_identity *static_identity,
|
||||
const u8 peer_public_key[NOISE_PUBLIC_KEY_LEN],
|
||||
const u8 peer_preshared_key[NOISE_SYMMETRIC_KEY_LEN],
|
||||
struct wg_peer *peer);
|
||||
void wg_noise_handshake_clear(struct noise_handshake *handshake);
|
||||
static inline void wg_noise_reset_last_sent_handshake(atomic64_t *handshake_ns)
|
||||
{
|
||||
atomic64_set(handshake_ns, ktime_get_coarse_boottime_ns() -
|
||||
(u64)(REKEY_TIMEOUT + 1) * NSEC_PER_SEC);
|
||||
}
|
||||
|
||||
void wg_noise_keypair_put(struct noise_keypair *keypair, bool unreference_now);
|
||||
struct noise_keypair *wg_noise_keypair_get(struct noise_keypair *keypair);
|
||||
void wg_noise_keypairs_clear(struct noise_keypairs *keypairs);
|
||||
bool wg_noise_received_with_keypair(struct noise_keypairs *keypairs,
|
||||
struct noise_keypair *received_keypair);
|
||||
void wg_noise_expire_current_peer_keypairs(struct wg_peer *peer);
|
||||
|
||||
void wg_noise_set_static_identity_private_key(
|
||||
struct noise_static_identity *static_identity,
|
||||
const u8 private_key[NOISE_PUBLIC_KEY_LEN]);
|
||||
void wg_noise_precompute_static_static(struct wg_peer *peer);
|
||||
|
||||
bool
|
||||
wg_noise_handshake_create_initiation(struct message_handshake_initiation *dst,
|
||||
struct noise_handshake *handshake);
|
||||
struct wg_peer *
|
||||
wg_noise_handshake_consume_initiation(struct message_handshake_initiation *src,
|
||||
struct wg_device *wg);
|
||||
|
||||
bool wg_noise_handshake_create_response(struct message_handshake_response *dst,
|
||||
struct noise_handshake *handshake);
|
||||
struct wg_peer *
|
||||
wg_noise_handshake_consume_response(struct message_handshake_response *src,
|
||||
struct wg_device *wg);
|
||||
|
||||
bool wg_noise_handshake_begin_session(struct noise_handshake *handshake,
|
||||
struct noise_keypairs *keypairs);
|
||||
|
||||
#endif /* _WG_NOISE_H */
|
||||
240
drivers/net/wireguard/peer.c
Normal file
240
drivers/net/wireguard/peer.c
Normal file
|
|
@ -0,0 +1,240 @@
|
|||
// SPDX-License-Identifier: GPL-2.0
|
||||
/*
|
||||
* Copyright (C) 2015-2019 Jason A. Donenfeld <Jason@zx2c4.com>. All Rights Reserved.
|
||||
*/
|
||||
|
||||
#include "peer.h"
|
||||
#include "device.h"
|
||||
#include "queueing.h"
|
||||
#include "timers.h"
|
||||
#include "peerlookup.h"
|
||||
#include "noise.h"
|
||||
|
||||
#include <linux/kref.h>
|
||||
#include <linux/lockdep.h>
|
||||
#include <linux/rcupdate.h>
|
||||
#include <linux/list.h>
|
||||
|
||||
static struct kmem_cache *peer_cache;
|
||||
static atomic64_t peer_counter = ATOMIC64_INIT(0);
|
||||
|
||||
struct wg_peer *wg_peer_create(struct wg_device *wg,
|
||||
const u8 public_key[NOISE_PUBLIC_KEY_LEN],
|
||||
const u8 preshared_key[NOISE_SYMMETRIC_KEY_LEN])
|
||||
{
|
||||
struct wg_peer *peer;
|
||||
int ret = -ENOMEM;
|
||||
|
||||
lockdep_assert_held(&wg->device_update_lock);
|
||||
|
||||
if (wg->num_peers >= MAX_PEERS_PER_DEVICE)
|
||||
return ERR_PTR(ret);
|
||||
|
||||
peer = kmem_cache_zalloc(peer_cache, GFP_KERNEL);
|
||||
if (unlikely(!peer))
|
||||
return ERR_PTR(ret);
|
||||
if (unlikely(dst_cache_init(&peer->endpoint_cache, GFP_KERNEL)))
|
||||
goto err;
|
||||
|
||||
peer->device = wg;
|
||||
wg_noise_handshake_init(&peer->handshake, &wg->static_identity,
|
||||
public_key, preshared_key, peer);
|
||||
peer->internal_id = atomic64_inc_return(&peer_counter);
|
||||
peer->serial_work_cpu = nr_cpumask_bits;
|
||||
wg_cookie_init(&peer->latest_cookie);
|
||||
wg_timers_init(peer);
|
||||
wg_cookie_checker_precompute_peer_keys(peer);
|
||||
spin_lock_init(&peer->keypairs.keypair_update_lock);
|
||||
INIT_WORK(&peer->transmit_handshake_work, wg_packet_handshake_send_worker);
|
||||
INIT_WORK(&peer->transmit_packet_work, wg_packet_tx_worker);
|
||||
wg_prev_queue_init(&peer->tx_queue);
|
||||
wg_prev_queue_init(&peer->rx_queue);
|
||||
rwlock_init(&peer->endpoint_lock);
|
||||
kref_init(&peer->refcount);
|
||||
skb_queue_head_init(&peer->staged_packet_queue);
|
||||
wg_noise_reset_last_sent_handshake(&peer->last_sent_handshake);
|
||||
set_bit(NAPI_STATE_NO_BUSY_POLL, &peer->napi.state);
|
||||
netif_napi_add(wg->dev, &peer->napi, wg_packet_rx_poll,
|
||||
NAPI_POLL_WEIGHT);
|
||||
napi_enable(&peer->napi);
|
||||
list_add_tail(&peer->peer_list, &wg->peer_list);
|
||||
INIT_LIST_HEAD(&peer->allowedips_list);
|
||||
wg_pubkey_hashtable_add(wg->peer_hashtable, peer);
|
||||
++wg->num_peers;
|
||||
pr_debug("%s: Peer %llu created\n", wg->dev->name, peer->internal_id);
|
||||
return peer;
|
||||
|
||||
err:
|
||||
kmem_cache_free(peer_cache, peer);
|
||||
return ERR_PTR(ret);
|
||||
}
|
||||
|
||||
struct wg_peer *wg_peer_get_maybe_zero(struct wg_peer *peer)
|
||||
{
|
||||
RCU_LOCKDEP_WARN(!rcu_read_lock_bh_held(),
|
||||
"Taking peer reference without holding the RCU read lock");
|
||||
if (unlikely(!peer || !kref_get_unless_zero(&peer->refcount)))
|
||||
return NULL;
|
||||
return peer;
|
||||
}
|
||||
|
||||
static void peer_make_dead(struct wg_peer *peer)
|
||||
{
|
||||
/* Remove from configuration-time lookup structures. */
|
||||
list_del_init(&peer->peer_list);
|
||||
wg_allowedips_remove_by_peer(&peer->device->peer_allowedips, peer,
|
||||
&peer->device->device_update_lock);
|
||||
wg_pubkey_hashtable_remove(peer->device->peer_hashtable, peer);
|
||||
|
||||
/* Mark as dead, so that we don't allow jumping contexts after. */
|
||||
WRITE_ONCE(peer->is_dead, true);
|
||||
|
||||
/* The caller must now synchronize_net() for this to take effect. */
|
||||
}
|
||||
|
||||
static void peer_remove_after_dead(struct wg_peer *peer)
|
||||
{
|
||||
WARN_ON(!peer->is_dead);
|
||||
|
||||
/* No more keypairs can be created for this peer, since is_dead protects
|
||||
* add_new_keypair, so we can now destroy existing ones.
|
||||
*/
|
||||
wg_noise_keypairs_clear(&peer->keypairs);
|
||||
|
||||
/* Destroy all ongoing timers that were in-flight at the beginning of
|
||||
* this function.
|
||||
*/
|
||||
wg_timers_stop(peer);
|
||||
|
||||
/* The transition between packet encryption/decryption queues isn't
|
||||
* guarded by is_dead, but each reference's life is strictly bounded by
|
||||
* two generations: once for parallel crypto and once for serial
|
||||
* ingestion, so we can simply flush twice, and be sure that we no
|
||||
* longer have references inside these queues.
|
||||
*/
|
||||
|
||||
/* a) For encrypt/decrypt. */
|
||||
flush_workqueue(peer->device->packet_crypt_wq);
|
||||
/* b.1) For send (but not receive, since that's napi). */
|
||||
flush_workqueue(peer->device->packet_crypt_wq);
|
||||
/* b.2.1) For receive (but not send, since that's wq). */
|
||||
napi_disable(&peer->napi);
|
||||
/* b.2.1) It's now safe to remove the napi struct, which must be done
|
||||
* here from process context.
|
||||
*/
|
||||
netif_napi_del(&peer->napi);
|
||||
|
||||
/* Ensure any workstructs we own (like transmit_handshake_work or
|
||||
* clear_peer_work) no longer are in use.
|
||||
*/
|
||||
flush_workqueue(peer->device->handshake_send_wq);
|
||||
|
||||
/* After the above flushes, a peer might still be active in a few
|
||||
* different contexts: 1) from xmit(), before hitting is_dead and
|
||||
* returning, 2) from wg_packet_consume_data(), before hitting is_dead
|
||||
* and returning, 3) from wg_receive_handshake_packet() after a point
|
||||
* where it has processed an incoming handshake packet, but where
|
||||
* all calls to pass it off to timers fails because of is_dead. We won't
|
||||
* have new references in (1) eventually, because we're removed from
|
||||
* allowedips; we won't have new references in (2) eventually, because
|
||||
* wg_index_hashtable_lookup will always return NULL, since we removed
|
||||
* all existing keypairs and no more can be created; we won't have new
|
||||
* references in (3) eventually, because we're removed from the pubkey
|
||||
* hash table, which allows for a maximum of one handshake response,
|
||||
* via the still-uncleared index hashtable entry, but not more than one,
|
||||
* and in wg_cookie_message_consume, the lookup eventually gets a peer
|
||||
* with a refcount of zero, so no new reference is taken.
|
||||
*/
|
||||
|
||||
--peer->device->num_peers;
|
||||
wg_peer_put(peer);
|
||||
}
|
||||
|
||||
/* We have a separate "remove" function make sure that all active places where
|
||||
* a peer is currently operating will eventually come to an end and not pass
|
||||
* their reference onto another context.
|
||||
*/
|
||||
void wg_peer_remove(struct wg_peer *peer)
|
||||
{
|
||||
if (unlikely(!peer))
|
||||
return;
|
||||
lockdep_assert_held(&peer->device->device_update_lock);
|
||||
|
||||
peer_make_dead(peer);
|
||||
synchronize_net();
|
||||
peer_remove_after_dead(peer);
|
||||
}
|
||||
|
||||
void wg_peer_remove_all(struct wg_device *wg)
|
||||
{
|
||||
struct wg_peer *peer, *temp;
|
||||
LIST_HEAD(dead_peers);
|
||||
|
||||
lockdep_assert_held(&wg->device_update_lock);
|
||||
|
||||
/* Avoid having to traverse individually for each one. */
|
||||
wg_allowedips_free(&wg->peer_allowedips, &wg->device_update_lock);
|
||||
|
||||
list_for_each_entry_safe(peer, temp, &wg->peer_list, peer_list) {
|
||||
peer_make_dead(peer);
|
||||
list_add_tail(&peer->peer_list, &dead_peers);
|
||||
}
|
||||
synchronize_net();
|
||||
list_for_each_entry_safe(peer, temp, &dead_peers, peer_list)
|
||||
peer_remove_after_dead(peer);
|
||||
}
|
||||
|
||||
static void rcu_release(struct rcu_head *rcu)
|
||||
{
|
||||
struct wg_peer *peer = container_of(rcu, struct wg_peer, rcu);
|
||||
|
||||
dst_cache_destroy(&peer->endpoint_cache);
|
||||
WARN_ON(wg_prev_queue_peek(&peer->tx_queue) || wg_prev_queue_peek(&peer->rx_queue));
|
||||
|
||||
/* The final zeroing takes care of clearing any remaining handshake key
|
||||
* material and other potentially sensitive information.
|
||||
*/
|
||||
memzero_explicit(peer, sizeof(*peer));
|
||||
kmem_cache_free(peer_cache, peer);
|
||||
}
|
||||
|
||||
static void kref_release(struct kref *refcount)
|
||||
{
|
||||
struct wg_peer *peer = container_of(refcount, struct wg_peer, refcount);
|
||||
|
||||
pr_debug("%s: Peer %llu (%pISpfsc) destroyed\n",
|
||||
peer->device->dev->name, peer->internal_id,
|
||||
&peer->endpoint.addr);
|
||||
|
||||
/* Remove ourself from dynamic runtime lookup structures, now that the
|
||||
* last reference is gone.
|
||||
*/
|
||||
wg_index_hashtable_remove(peer->device->index_hashtable,
|
||||
&peer->handshake.entry);
|
||||
|
||||
/* Remove any lingering packets that didn't have a chance to be
|
||||
* transmitted.
|
||||
*/
|
||||
wg_packet_purge_staged_packets(peer);
|
||||
|
||||
/* Free the memory used. */
|
||||
call_rcu(&peer->rcu, rcu_release);
|
||||
}
|
||||
|
||||
void wg_peer_put(struct wg_peer *peer)
|
||||
{
|
||||
if (unlikely(!peer))
|
||||
return;
|
||||
kref_put(&peer->refcount, kref_release);
|
||||
}
|
||||
|
||||
int __init wg_peer_init(void)
|
||||
{
|
||||
peer_cache = KMEM_CACHE(wg_peer, 0);
|
||||
return peer_cache ? 0 : -ENOMEM;
|
||||
}
|
||||
|
||||
void wg_peer_uninit(void)
|
||||
{
|
||||
kmem_cache_destroy(peer_cache);
|
||||
}
|
||||
86
drivers/net/wireguard/peer.h
Normal file
86
drivers/net/wireguard/peer.h
Normal file
|
|
@ -0,0 +1,86 @@
|
|||
/* SPDX-License-Identifier: GPL-2.0 */
|
||||
/*
|
||||
* Copyright (C) 2015-2019 Jason A. Donenfeld <Jason@zx2c4.com>. All Rights Reserved.
|
||||
*/
|
||||
|
||||
#ifndef _WG_PEER_H
|
||||
#define _WG_PEER_H
|
||||
|
||||
#include "device.h"
|
||||
#include "noise.h"
|
||||
#include "cookie.h"
|
||||
|
||||
#include <linux/types.h>
|
||||
#include <linux/netfilter.h>
|
||||
#include <linux/spinlock.h>
|
||||
#include <linux/kref.h>
|
||||
#include <net/dst_cache.h>
|
||||
|
||||
struct wg_device;
|
||||
|
||||
struct endpoint {
|
||||
union {
|
||||
struct sockaddr addr;
|
||||
struct sockaddr_in addr4;
|
||||
struct sockaddr_in6 addr6;
|
||||
};
|
||||
union {
|
||||
struct {
|
||||
struct in_addr src4;
|
||||
/* Essentially the same as addr6->scope_id */
|
||||
int src_if4;
|
||||
};
|
||||
struct in6_addr src6;
|
||||
};
|
||||
};
|
||||
|
||||
struct wg_peer {
|
||||
struct wg_device *device;
|
||||
struct prev_queue tx_queue, rx_queue;
|
||||
struct sk_buff_head staged_packet_queue;
|
||||
int serial_work_cpu;
|
||||
struct noise_keypairs keypairs;
|
||||
struct endpoint endpoint;
|
||||
struct dst_cache endpoint_cache;
|
||||
rwlock_t endpoint_lock;
|
||||
struct noise_handshake handshake;
|
||||
atomic64_t last_sent_handshake;
|
||||
struct work_struct transmit_handshake_work, clear_peer_work, transmit_packet_work;
|
||||
struct cookie latest_cookie;
|
||||
struct hlist_node pubkey_hash;
|
||||
u64 rx_bytes, tx_bytes;
|
||||
struct timer_list timer_retransmit_handshake, timer_send_keepalive;
|
||||
struct timer_list timer_new_handshake, timer_zero_key_material;
|
||||
struct timer_list timer_persistent_keepalive;
|
||||
unsigned int timer_handshake_attempts;
|
||||
u16 persistent_keepalive_interval;
|
||||
bool timer_need_another_keepalive;
|
||||
bool sent_lastminute_handshake;
|
||||
struct timespec64 walltime_last_handshake;
|
||||
struct kref refcount;
|
||||
struct rcu_head rcu;
|
||||
struct list_head peer_list;
|
||||
struct list_head allowedips_list;
|
||||
u64 internal_id;
|
||||
struct napi_struct napi;
|
||||
bool is_dead;
|
||||
};
|
||||
|
||||
struct wg_peer *wg_peer_create(struct wg_device *wg,
|
||||
const u8 public_key[NOISE_PUBLIC_KEY_LEN],
|
||||
const u8 preshared_key[NOISE_SYMMETRIC_KEY_LEN]);
|
||||
|
||||
struct wg_peer *__must_check wg_peer_get_maybe_zero(struct wg_peer *peer);
|
||||
static inline struct wg_peer *wg_peer_get(struct wg_peer *peer)
|
||||
{
|
||||
kref_get(&peer->refcount);
|
||||
return peer;
|
||||
}
|
||||
void wg_peer_put(struct wg_peer *peer);
|
||||
void wg_peer_remove(struct wg_peer *peer);
|
||||
void wg_peer_remove_all(struct wg_device *wg);
|
||||
|
||||
int wg_peer_init(void);
|
||||
void wg_peer_uninit(void);
|
||||
|
||||
#endif /* _WG_PEER_H */
|
||||
226
drivers/net/wireguard/peerlookup.c
Normal file
226
drivers/net/wireguard/peerlookup.c
Normal file
|
|
@ -0,0 +1,226 @@
|
|||
// SPDX-License-Identifier: GPL-2.0
|
||||
/*
|
||||
* Copyright (C) 2015-2019 Jason A. Donenfeld <Jason@zx2c4.com>. All Rights Reserved.
|
||||
*/
|
||||
|
||||
#include "peerlookup.h"
|
||||
#include "peer.h"
|
||||
#include "noise.h"
|
||||
|
||||
static struct hlist_head *pubkey_bucket(struct pubkey_hashtable *table,
|
||||
const u8 pubkey[NOISE_PUBLIC_KEY_LEN])
|
||||
{
|
||||
/* siphash gives us a secure 64bit number based on a random key. Since
|
||||
* the bits are uniformly distributed, we can then mask off to get the
|
||||
* bits we need.
|
||||
*/
|
||||
const u64 hash = siphash(pubkey, NOISE_PUBLIC_KEY_LEN, &table->key);
|
||||
|
||||
return &table->hashtable[hash & (HASH_SIZE(table->hashtable) - 1)];
|
||||
}
|
||||
|
||||
struct pubkey_hashtable *wg_pubkey_hashtable_alloc(void)
|
||||
{
|
||||
struct pubkey_hashtable *table = kvmalloc(sizeof(*table), GFP_KERNEL);
|
||||
|
||||
if (!table)
|
||||
return NULL;
|
||||
|
||||
get_random_bytes(&table->key, sizeof(table->key));
|
||||
hash_init(table->hashtable);
|
||||
mutex_init(&table->lock);
|
||||
return table;
|
||||
}
|
||||
|
||||
void wg_pubkey_hashtable_add(struct pubkey_hashtable *table,
|
||||
struct wg_peer *peer)
|
||||
{
|
||||
mutex_lock(&table->lock);
|
||||
hlist_add_head_rcu(&peer->pubkey_hash,
|
||||
pubkey_bucket(table, peer->handshake.remote_static));
|
||||
mutex_unlock(&table->lock);
|
||||
}
|
||||
|
||||
void wg_pubkey_hashtable_remove(struct pubkey_hashtable *table,
|
||||
struct wg_peer *peer)
|
||||
{
|
||||
mutex_lock(&table->lock);
|
||||
hlist_del_init_rcu(&peer->pubkey_hash);
|
||||
mutex_unlock(&table->lock);
|
||||
}
|
||||
|
||||
/* Returns a strong reference to a peer */
|
||||
struct wg_peer *
|
||||
wg_pubkey_hashtable_lookup(struct pubkey_hashtable *table,
|
||||
const u8 pubkey[NOISE_PUBLIC_KEY_LEN])
|
||||
{
|
||||
struct wg_peer *iter_peer, *peer = NULL;
|
||||
|
||||
rcu_read_lock_bh();
|
||||
hlist_for_each_entry_rcu_bh(iter_peer, pubkey_bucket(table, pubkey),
|
||||
pubkey_hash) {
|
||||
if (!memcmp(pubkey, iter_peer->handshake.remote_static,
|
||||
NOISE_PUBLIC_KEY_LEN)) {
|
||||
peer = iter_peer;
|
||||
break;
|
||||
}
|
||||
}
|
||||
peer = wg_peer_get_maybe_zero(peer);
|
||||
rcu_read_unlock_bh();
|
||||
return peer;
|
||||
}
|
||||
|
||||
static struct hlist_head *index_bucket(struct index_hashtable *table,
|
||||
const __le32 index)
|
||||
{
|
||||
/* Since the indices are random and thus all bits are uniformly
|
||||
* distributed, we can find its bucket simply by masking.
|
||||
*/
|
||||
return &table->hashtable[(__force u32)index &
|
||||
(HASH_SIZE(table->hashtable) - 1)];
|
||||
}
|
||||
|
||||
struct index_hashtable *wg_index_hashtable_alloc(void)
|
||||
{
|
||||
struct index_hashtable *table = kvmalloc(sizeof(*table), GFP_KERNEL);
|
||||
|
||||
if (!table)
|
||||
return NULL;
|
||||
|
||||
hash_init(table->hashtable);
|
||||
spin_lock_init(&table->lock);
|
||||
return table;
|
||||
}
|
||||
|
||||
/* At the moment, we limit ourselves to 2^20 total peers, which generally might
|
||||
* amount to 2^20*3 items in this hashtable. The algorithm below works by
|
||||
* picking a random number and testing it. We can see that these limits mean we
|
||||
* usually succeed pretty quickly:
|
||||
*
|
||||
* >>> def calculation(tries, size):
|
||||
* ... return (size / 2**32)**(tries - 1) * (1 - (size / 2**32))
|
||||
* ...
|
||||
* >>> calculation(1, 2**20 * 3)
|
||||
* 0.999267578125
|
||||
* >>> calculation(2, 2**20 * 3)
|
||||
* 0.0007318854331970215
|
||||
* >>> calculation(3, 2**20 * 3)
|
||||
* 5.360489012673497e-07
|
||||
* >>> calculation(4, 2**20 * 3)
|
||||
* 3.9261394135792216e-10
|
||||
*
|
||||
* At the moment, we don't do any masking, so this algorithm isn't exactly
|
||||
* constant time in either the random guessing or in the hash list lookup. We
|
||||
* could require a minimum of 3 tries, which would successfully mask the
|
||||
* guessing. this would not, however, help with the growing hash lengths, which
|
||||
* is another thing to consider moving forward.
|
||||
*/
|
||||
|
||||
__le32 wg_index_hashtable_insert(struct index_hashtable *table,
|
||||
struct index_hashtable_entry *entry)
|
||||
{
|
||||
struct index_hashtable_entry *existing_entry;
|
||||
|
||||
spin_lock_bh(&table->lock);
|
||||
hlist_del_init_rcu(&entry->index_hash);
|
||||
spin_unlock_bh(&table->lock);
|
||||
|
||||
rcu_read_lock_bh();
|
||||
|
||||
search_unused_slot:
|
||||
/* First we try to find an unused slot, randomly, while unlocked. */
|
||||
entry->index = (__force __le32)get_random_u32();
|
||||
hlist_for_each_entry_rcu_bh(existing_entry,
|
||||
index_bucket(table, entry->index),
|
||||
index_hash) {
|
||||
if (existing_entry->index == entry->index)
|
||||
/* If it's already in use, we continue searching. */
|
||||
goto search_unused_slot;
|
||||
}
|
||||
|
||||
/* Once we've found an unused slot, we lock it, and then double-check
|
||||
* that nobody else stole it from us.
|
||||
*/
|
||||
spin_lock_bh(&table->lock);
|
||||
hlist_for_each_entry_rcu_bh(existing_entry,
|
||||
index_bucket(table, entry->index),
|
||||
index_hash) {
|
||||
if (existing_entry->index == entry->index) {
|
||||
spin_unlock_bh(&table->lock);
|
||||
/* If it was stolen, we start over. */
|
||||
goto search_unused_slot;
|
||||
}
|
||||
}
|
||||
/* Otherwise, we know we have it exclusively (since we're locked),
|
||||
* so we insert.
|
||||
*/
|
||||
hlist_add_head_rcu(&entry->index_hash,
|
||||
index_bucket(table, entry->index));
|
||||
spin_unlock_bh(&table->lock);
|
||||
|
||||
rcu_read_unlock_bh();
|
||||
|
||||
return entry->index;
|
||||
}
|
||||
|
||||
bool wg_index_hashtable_replace(struct index_hashtable *table,
|
||||
struct index_hashtable_entry *old,
|
||||
struct index_hashtable_entry *new)
|
||||
{
|
||||
bool ret;
|
||||
|
||||
spin_lock_bh(&table->lock);
|
||||
ret = !hlist_unhashed(&old->index_hash);
|
||||
if (unlikely(!ret))
|
||||
goto out;
|
||||
|
||||
new->index = old->index;
|
||||
hlist_replace_rcu(&old->index_hash, &new->index_hash);
|
||||
|
||||
/* Calling init here NULLs out index_hash, and in fact after this
|
||||
* function returns, it's theoretically possible for this to get
|
||||
* reinserted elsewhere. That means the RCU lookup below might either
|
||||
* terminate early or jump between buckets, in which case the packet
|
||||
* simply gets dropped, which isn't terrible.
|
||||
*/
|
||||
INIT_HLIST_NODE(&old->index_hash);
|
||||
out:
|
||||
spin_unlock_bh(&table->lock);
|
||||
return ret;
|
||||
}
|
||||
|
||||
void wg_index_hashtable_remove(struct index_hashtable *table,
|
||||
struct index_hashtable_entry *entry)
|
||||
{
|
||||
spin_lock_bh(&table->lock);
|
||||
hlist_del_init_rcu(&entry->index_hash);
|
||||
spin_unlock_bh(&table->lock);
|
||||
}
|
||||
|
||||
/* Returns a strong reference to a entry->peer */
|
||||
struct index_hashtable_entry *
|
||||
wg_index_hashtable_lookup(struct index_hashtable *table,
|
||||
const enum index_hashtable_type type_mask,
|
||||
const __le32 index, struct wg_peer **peer)
|
||||
{
|
||||
struct index_hashtable_entry *iter_entry, *entry = NULL;
|
||||
|
||||
rcu_read_lock_bh();
|
||||
hlist_for_each_entry_rcu_bh(iter_entry, index_bucket(table, index),
|
||||
index_hash) {
|
||||
if (iter_entry->index == index) {
|
||||
if (likely(iter_entry->type & type_mask))
|
||||
entry = iter_entry;
|
||||
break;
|
||||
}
|
||||
}
|
||||
if (likely(entry)) {
|
||||
entry->peer = wg_peer_get_maybe_zero(entry->peer);
|
||||
if (likely(entry->peer))
|
||||
*peer = entry->peer;
|
||||
else
|
||||
entry = NULL;
|
||||
}
|
||||
rcu_read_unlock_bh();
|
||||
return entry;
|
||||
}
|
||||
64
drivers/net/wireguard/peerlookup.h
Normal file
64
drivers/net/wireguard/peerlookup.h
Normal file
|
|
@ -0,0 +1,64 @@
|
|||
/* SPDX-License-Identifier: GPL-2.0 */
|
||||
/*
|
||||
* Copyright (C) 2015-2019 Jason A. Donenfeld <Jason@zx2c4.com>. All Rights Reserved.
|
||||
*/
|
||||
|
||||
#ifndef _WG_PEERLOOKUP_H
|
||||
#define _WG_PEERLOOKUP_H
|
||||
|
||||
#include "messages.h"
|
||||
|
||||
#include <linux/hashtable.h>
|
||||
#include <linux/mutex.h>
|
||||
#include <linux/siphash.h>
|
||||
|
||||
struct wg_peer;
|
||||
|
||||
struct pubkey_hashtable {
|
||||
/* TODO: move to rhashtable */
|
||||
DECLARE_HASHTABLE(hashtable, 11);
|
||||
siphash_key_t key;
|
||||
struct mutex lock;
|
||||
};
|
||||
|
||||
struct pubkey_hashtable *wg_pubkey_hashtable_alloc(void);
|
||||
void wg_pubkey_hashtable_add(struct pubkey_hashtable *table,
|
||||
struct wg_peer *peer);
|
||||
void wg_pubkey_hashtable_remove(struct pubkey_hashtable *table,
|
||||
struct wg_peer *peer);
|
||||
struct wg_peer *
|
||||
wg_pubkey_hashtable_lookup(struct pubkey_hashtable *table,
|
||||
const u8 pubkey[NOISE_PUBLIC_KEY_LEN]);
|
||||
|
||||
struct index_hashtable {
|
||||
/* TODO: move to rhashtable */
|
||||
DECLARE_HASHTABLE(hashtable, 13);
|
||||
spinlock_t lock;
|
||||
};
|
||||
|
||||
enum index_hashtable_type {
|
||||
INDEX_HASHTABLE_HANDSHAKE = 1U << 0,
|
||||
INDEX_HASHTABLE_KEYPAIR = 1U << 1
|
||||
};
|
||||
|
||||
struct index_hashtable_entry {
|
||||
struct wg_peer *peer;
|
||||
struct hlist_node index_hash;
|
||||
enum index_hashtable_type type;
|
||||
__le32 index;
|
||||
};
|
||||
|
||||
struct index_hashtable *wg_index_hashtable_alloc(void);
|
||||
__le32 wg_index_hashtable_insert(struct index_hashtable *table,
|
||||
struct index_hashtable_entry *entry);
|
||||
bool wg_index_hashtable_replace(struct index_hashtable *table,
|
||||
struct index_hashtable_entry *old,
|
||||
struct index_hashtable_entry *new);
|
||||
void wg_index_hashtable_remove(struct index_hashtable *table,
|
||||
struct index_hashtable_entry *entry);
|
||||
struct index_hashtable_entry *
|
||||
wg_index_hashtable_lookup(struct index_hashtable *table,
|
||||
const enum index_hashtable_type type_mask,
|
||||
const __le32 index, struct wg_peer **peer);
|
||||
|
||||
#endif /* _WG_PEERLOOKUP_H */
|
||||
109
drivers/net/wireguard/queueing.c
Normal file
109
drivers/net/wireguard/queueing.c
Normal file
|
|
@ -0,0 +1,109 @@
|
|||
// SPDX-License-Identifier: GPL-2.0
|
||||
/*
|
||||
* Copyright (C) 2015-2019 Jason A. Donenfeld <Jason@zx2c4.com>. All Rights Reserved.
|
||||
*/
|
||||
|
||||
#include "queueing.h"
|
||||
#include <linux/skb_array.h>
|
||||
|
||||
struct multicore_worker __percpu *
|
||||
wg_packet_percpu_multicore_worker_alloc(work_func_t function, void *ptr)
|
||||
{
|
||||
int cpu;
|
||||
struct multicore_worker __percpu *worker = alloc_percpu(struct multicore_worker);
|
||||
|
||||
if (!worker)
|
||||
return NULL;
|
||||
|
||||
for_each_possible_cpu(cpu) {
|
||||
per_cpu_ptr(worker, cpu)->ptr = ptr;
|
||||
INIT_WORK(&per_cpu_ptr(worker, cpu)->work, function);
|
||||
}
|
||||
return worker;
|
||||
}
|
||||
|
||||
int wg_packet_queue_init(struct crypt_queue *queue, work_func_t function,
|
||||
unsigned int len)
|
||||
{
|
||||
int ret;
|
||||
|
||||
memset(queue, 0, sizeof(*queue));
|
||||
queue->last_cpu = -1;
|
||||
ret = ptr_ring_init(&queue->ring, len, GFP_KERNEL);
|
||||
if (ret)
|
||||
return ret;
|
||||
queue->worker = wg_packet_percpu_multicore_worker_alloc(function, queue);
|
||||
if (!queue->worker) {
|
||||
ptr_ring_cleanup(&queue->ring, NULL);
|
||||
return -ENOMEM;
|
||||
}
|
||||
return 0;
|
||||
}
|
||||
|
||||
void wg_packet_queue_free(struct crypt_queue *queue, bool purge)
|
||||
{
|
||||
free_percpu(queue->worker);
|
||||
WARN_ON(!purge && !__ptr_ring_empty(&queue->ring));
|
||||
ptr_ring_cleanup(&queue->ring, purge ? __skb_array_destroy_skb : NULL);
|
||||
}
|
||||
|
||||
#define NEXT(skb) ((skb)->prev)
|
||||
#define STUB(queue) ((struct sk_buff *)&queue->empty)
|
||||
|
||||
void wg_prev_queue_init(struct prev_queue *queue)
|
||||
{
|
||||
NEXT(STUB(queue)) = NULL;
|
||||
queue->head = queue->tail = STUB(queue);
|
||||
queue->peeked = NULL;
|
||||
atomic_set(&queue->count, 0);
|
||||
BUILD_BUG_ON(
|
||||
offsetof(struct sk_buff, next) != offsetof(struct prev_queue, empty.next) -
|
||||
offsetof(struct prev_queue, empty) ||
|
||||
offsetof(struct sk_buff, prev) != offsetof(struct prev_queue, empty.prev) -
|
||||
offsetof(struct prev_queue, empty));
|
||||
}
|
||||
|
||||
static void __wg_prev_queue_enqueue(struct prev_queue *queue, struct sk_buff *skb)
|
||||
{
|
||||
WRITE_ONCE(NEXT(skb), NULL);
|
||||
WRITE_ONCE(NEXT(xchg_release(&queue->head, skb)), skb);
|
||||
}
|
||||
|
||||
bool wg_prev_queue_enqueue(struct prev_queue *queue, struct sk_buff *skb)
|
||||
{
|
||||
if (!atomic_add_unless(&queue->count, 1, MAX_QUEUED_PACKETS))
|
||||
return false;
|
||||
__wg_prev_queue_enqueue(queue, skb);
|
||||
return true;
|
||||
}
|
||||
|
||||
struct sk_buff *wg_prev_queue_dequeue(struct prev_queue *queue)
|
||||
{
|
||||
struct sk_buff *tail = queue->tail, *next = smp_load_acquire(&NEXT(tail));
|
||||
|
||||
if (tail == STUB(queue)) {
|
||||
if (!next)
|
||||
return NULL;
|
||||
queue->tail = next;
|
||||
tail = next;
|
||||
next = smp_load_acquire(&NEXT(next));
|
||||
}
|
||||
if (next) {
|
||||
queue->tail = next;
|
||||
atomic_dec(&queue->count);
|
||||
return tail;
|
||||
}
|
||||
if (tail != READ_ONCE(queue->head))
|
||||
return NULL;
|
||||
__wg_prev_queue_enqueue(queue, STUB(queue));
|
||||
next = smp_load_acquire(&NEXT(tail));
|
||||
if (next) {
|
||||
queue->tail = next;
|
||||
atomic_dec(&queue->count);
|
||||
return tail;
|
||||
}
|
||||
return NULL;
|
||||
}
|
||||
|
||||
#undef NEXT
|
||||
#undef STUB
|
||||
222
drivers/net/wireguard/queueing.h
Normal file
222
drivers/net/wireguard/queueing.h
Normal file
|
|
@ -0,0 +1,222 @@
|
|||
/* SPDX-License-Identifier: GPL-2.0 */
|
||||
/*
|
||||
* Copyright (C) 2015-2019 Jason A. Donenfeld <Jason@zx2c4.com>. All Rights Reserved.
|
||||
*/
|
||||
|
||||
#ifndef _WG_QUEUEING_H
|
||||
#define _WG_QUEUEING_H
|
||||
|
||||
#include "peer.h"
|
||||
#include <linux/types.h>
|
||||
#include <linux/skbuff.h>
|
||||
#include <linux/ip.h>
|
||||
#include <linux/ipv6.h>
|
||||
#include <net/ip_tunnels.h>
|
||||
|
||||
struct wg_device;
|
||||
struct wg_peer;
|
||||
struct multicore_worker;
|
||||
struct crypt_queue;
|
||||
struct prev_queue;
|
||||
struct sk_buff;
|
||||
|
||||
/* queueing.c APIs: */
|
||||
int wg_packet_queue_init(struct crypt_queue *queue, work_func_t function,
|
||||
unsigned int len);
|
||||
void wg_packet_queue_free(struct crypt_queue *queue, bool purge);
|
||||
struct multicore_worker __percpu *
|
||||
wg_packet_percpu_multicore_worker_alloc(work_func_t function, void *ptr);
|
||||
|
||||
/* receive.c APIs: */
|
||||
void wg_packet_receive(struct wg_device *wg, struct sk_buff *skb);
|
||||
void wg_packet_handshake_receive_worker(struct work_struct *work);
|
||||
/* NAPI poll function: */
|
||||
int wg_packet_rx_poll(struct napi_struct *napi, int budget);
|
||||
/* Workqueue worker: */
|
||||
void wg_packet_decrypt_worker(struct work_struct *work);
|
||||
|
||||
/* send.c APIs: */
|
||||
void wg_packet_send_queued_handshake_initiation(struct wg_peer *peer,
|
||||
bool is_retry);
|
||||
void wg_packet_send_handshake_response(struct wg_peer *peer);
|
||||
void wg_packet_send_handshake_cookie(struct wg_device *wg,
|
||||
struct sk_buff *initiating_skb,
|
||||
__le32 sender_index);
|
||||
void wg_packet_send_keepalive(struct wg_peer *peer);
|
||||
void wg_packet_purge_staged_packets(struct wg_peer *peer);
|
||||
void wg_packet_send_staged_packets(struct wg_peer *peer);
|
||||
/* Workqueue workers: */
|
||||
void wg_packet_handshake_send_worker(struct work_struct *work);
|
||||
void wg_packet_tx_worker(struct work_struct *work);
|
||||
void wg_packet_encrypt_worker(struct work_struct *work);
|
||||
|
||||
enum packet_state {
|
||||
PACKET_STATE_UNCRYPTED,
|
||||
PACKET_STATE_CRYPTED,
|
||||
PACKET_STATE_DEAD
|
||||
};
|
||||
|
||||
struct packet_cb {
|
||||
u64 nonce;
|
||||
struct noise_keypair *keypair;
|
||||
atomic_t state;
|
||||
u32 mtu;
|
||||
u8 ds;
|
||||
};
|
||||
|
||||
#define PACKET_CB(skb) ((struct packet_cb *)((skb)->cb))
|
||||
#define PACKET_PEER(skb) (PACKET_CB(skb)->keypair->entry.peer)
|
||||
|
||||
static inline bool wg_check_packet_protocol(struct sk_buff *skb)
|
||||
{
|
||||
__be16 real_protocol = ip_tunnel_parse_protocol(skb);
|
||||
return real_protocol && skb->protocol == real_protocol;
|
||||
}
|
||||
|
||||
static inline void wg_reset_packet(struct sk_buff *skb, bool encapsulating)
|
||||
{
|
||||
u8 l4_hash = skb->l4_hash;
|
||||
u8 sw_hash = skb->sw_hash;
|
||||
u32 hash = skb->hash;
|
||||
skb_scrub_packet(skb, true);
|
||||
memset(&skb->headers_start, 0,
|
||||
offsetof(struct sk_buff, headers_end) -
|
||||
offsetof(struct sk_buff, headers_start));
|
||||
|
||||
/* ANDROID:
|
||||
* Due to attempts to keep the ABI stable for struct sk_buff, the new
|
||||
* fields were incorrectly added _AFTER_ the headers_end field, which
|
||||
* requires that we manually copy the fields here from the old to the
|
||||
* new one.
|
||||
* Be sure to add any new field that is added in the
|
||||
* ANDROID_KABI_REPLACE() macros below here as well.
|
||||
*/
|
||||
skb->scm_io_uring = 0;
|
||||
|
||||
if (encapsulating) {
|
||||
skb->l4_hash = l4_hash;
|
||||
skb->sw_hash = sw_hash;
|
||||
skb->hash = hash;
|
||||
}
|
||||
skb->queue_mapping = 0;
|
||||
skb->nohdr = 0;
|
||||
skb->peeked = 0;
|
||||
skb->mac_len = 0;
|
||||
skb->dev = NULL;
|
||||
#ifdef CONFIG_NET_SCHED
|
||||
skb->tc_index = 0;
|
||||
#endif
|
||||
skb_reset_redirect(skb);
|
||||
skb->hdr_len = skb_headroom(skb);
|
||||
skb_reset_mac_header(skb);
|
||||
skb_reset_network_header(skb);
|
||||
skb_reset_transport_header(skb);
|
||||
skb_probe_transport_header(skb);
|
||||
skb_reset_inner_headers(skb);
|
||||
}
|
||||
|
||||
static inline int wg_cpumask_choose_online(int *stored_cpu, unsigned int id)
|
||||
{
|
||||
unsigned int cpu = *stored_cpu, cpu_index, i;
|
||||
|
||||
if (unlikely(cpu == nr_cpumask_bits ||
|
||||
!cpumask_test_cpu(cpu, cpu_online_mask))) {
|
||||
cpu_index = id % cpumask_weight(cpu_online_mask);
|
||||
cpu = cpumask_first(cpu_online_mask);
|
||||
for (i = 0; i < cpu_index; ++i)
|
||||
cpu = cpumask_next(cpu, cpu_online_mask);
|
||||
*stored_cpu = cpu;
|
||||
}
|
||||
return cpu;
|
||||
}
|
||||
|
||||
/* This function is racy, in the sense that it's called while last_cpu is
|
||||
* unlocked, so it could return the same CPU twice. Adding locking or using
|
||||
* atomic sequence numbers is slower though, and the consequences of racing are
|
||||
* harmless, so live with it.
|
||||
*/
|
||||
static inline int wg_cpumask_next_online(int *last_cpu)
|
||||
{
|
||||
int cpu = cpumask_next(READ_ONCE(*last_cpu), cpu_online_mask);
|
||||
if (cpu >= nr_cpu_ids)
|
||||
cpu = cpumask_first(cpu_online_mask);
|
||||
WRITE_ONCE(*last_cpu, cpu);
|
||||
return cpu;
|
||||
}
|
||||
|
||||
void wg_prev_queue_init(struct prev_queue *queue);
|
||||
|
||||
/* Multi producer */
|
||||
bool wg_prev_queue_enqueue(struct prev_queue *queue, struct sk_buff *skb);
|
||||
|
||||
/* Single consumer */
|
||||
struct sk_buff *wg_prev_queue_dequeue(struct prev_queue *queue);
|
||||
|
||||
/* Single consumer */
|
||||
static inline struct sk_buff *wg_prev_queue_peek(struct prev_queue *queue)
|
||||
{
|
||||
if (queue->peeked)
|
||||
return queue->peeked;
|
||||
queue->peeked = wg_prev_queue_dequeue(queue);
|
||||
return queue->peeked;
|
||||
}
|
||||
|
||||
/* Single consumer */
|
||||
static inline void wg_prev_queue_drop_peeked(struct prev_queue *queue)
|
||||
{
|
||||
queue->peeked = NULL;
|
||||
}
|
||||
|
||||
static inline int wg_queue_enqueue_per_device_and_peer(
|
||||
struct crypt_queue *device_queue, struct prev_queue *peer_queue,
|
||||
struct sk_buff *skb, struct workqueue_struct *wq)
|
||||
{
|
||||
int cpu;
|
||||
|
||||
atomic_set_release(&PACKET_CB(skb)->state, PACKET_STATE_UNCRYPTED);
|
||||
/* We first queue this up for the peer ingestion, but the consumer
|
||||
* will wait for the state to change to CRYPTED or DEAD before.
|
||||
*/
|
||||
if (unlikely(!wg_prev_queue_enqueue(peer_queue, skb)))
|
||||
return -ENOSPC;
|
||||
|
||||
/* Then we queue it up in the device queue, which consumes the
|
||||
* packet as soon as it can.
|
||||
*/
|
||||
cpu = wg_cpumask_next_online(&device_queue->last_cpu);
|
||||
if (unlikely(ptr_ring_produce_bh(&device_queue->ring, skb)))
|
||||
return -EPIPE;
|
||||
queue_work_on(cpu, wq, &per_cpu_ptr(device_queue->worker, cpu)->work);
|
||||
return 0;
|
||||
}
|
||||
|
||||
static inline void wg_queue_enqueue_per_peer_tx(struct sk_buff *skb, enum packet_state state)
|
||||
{
|
||||
/* We take a reference, because as soon as we call atomic_set, the
|
||||
* peer can be freed from below us.
|
||||
*/
|
||||
struct wg_peer *peer = wg_peer_get(PACKET_PEER(skb));
|
||||
|
||||
atomic_set_release(&PACKET_CB(skb)->state, state);
|
||||
queue_work_on(wg_cpumask_choose_online(&peer->serial_work_cpu, peer->internal_id),
|
||||
peer->device->packet_crypt_wq, &peer->transmit_packet_work);
|
||||
wg_peer_put(peer);
|
||||
}
|
||||
|
||||
static inline void wg_queue_enqueue_per_peer_rx(struct sk_buff *skb, enum packet_state state)
|
||||
{
|
||||
/* We take a reference, because as soon as we call atomic_set, the
|
||||
* peer can be freed from below us.
|
||||
*/
|
||||
struct wg_peer *peer = wg_peer_get(PACKET_PEER(skb));
|
||||
|
||||
atomic_set_release(&PACKET_CB(skb)->state, state);
|
||||
napi_schedule(&peer->napi);
|
||||
wg_peer_put(peer);
|
||||
}
|
||||
|
||||
#ifdef DEBUG
|
||||
bool wg_packet_counter_selftest(void);
|
||||
#endif
|
||||
|
||||
#endif /* _WG_QUEUEING_H */
|
||||
223
drivers/net/wireguard/ratelimiter.c
Normal file
223
drivers/net/wireguard/ratelimiter.c
Normal file
|
|
@ -0,0 +1,223 @@
|
|||
// SPDX-License-Identifier: GPL-2.0
|
||||
/*
|
||||
* Copyright (C) 2015-2019 Jason A. Donenfeld <Jason@zx2c4.com>. All Rights Reserved.
|
||||
*/
|
||||
|
||||
#include "ratelimiter.h"
|
||||
#include <linux/siphash.h>
|
||||
#include <linux/mm.h>
|
||||
#include <linux/slab.h>
|
||||
#include <net/ip.h>
|
||||
|
||||
static struct kmem_cache *entry_cache;
|
||||
static hsiphash_key_t key;
|
||||
static spinlock_t table_lock = __SPIN_LOCK_UNLOCKED("ratelimiter_table_lock");
|
||||
static DEFINE_MUTEX(init_lock);
|
||||
static u64 init_refcnt; /* Protected by init_lock, hence not atomic. */
|
||||
static atomic_t total_entries = ATOMIC_INIT(0);
|
||||
static unsigned int max_entries, table_size;
|
||||
static void wg_ratelimiter_gc_entries(struct work_struct *);
|
||||
static DECLARE_DEFERRABLE_WORK(gc_work, wg_ratelimiter_gc_entries);
|
||||
static struct hlist_head *table_v4;
|
||||
#if IS_ENABLED(CONFIG_IPV6)
|
||||
static struct hlist_head *table_v6;
|
||||
#endif
|
||||
|
||||
struct ratelimiter_entry {
|
||||
u64 last_time_ns, tokens, ip;
|
||||
void *net;
|
||||
spinlock_t lock;
|
||||
struct hlist_node hash;
|
||||
struct rcu_head rcu;
|
||||
};
|
||||
|
||||
enum {
|
||||
PACKETS_PER_SECOND = 20,
|
||||
PACKETS_BURSTABLE = 5,
|
||||
PACKET_COST = NSEC_PER_SEC / PACKETS_PER_SECOND,
|
||||
TOKEN_MAX = PACKET_COST * PACKETS_BURSTABLE
|
||||
};
|
||||
|
||||
static void entry_free(struct rcu_head *rcu)
|
||||
{
|
||||
kmem_cache_free(entry_cache,
|
||||
container_of(rcu, struct ratelimiter_entry, rcu));
|
||||
atomic_dec(&total_entries);
|
||||
}
|
||||
|
||||
static void entry_uninit(struct ratelimiter_entry *entry)
|
||||
{
|
||||
hlist_del_rcu(&entry->hash);
|
||||
call_rcu(&entry->rcu, entry_free);
|
||||
}
|
||||
|
||||
/* Calling this function with a NULL work uninits all entries. */
|
||||
static void wg_ratelimiter_gc_entries(struct work_struct *work)
|
||||
{
|
||||
const u64 now = ktime_get_coarse_boottime_ns();
|
||||
struct ratelimiter_entry *entry;
|
||||
struct hlist_node *temp;
|
||||
unsigned int i;
|
||||
|
||||
for (i = 0; i < table_size; ++i) {
|
||||
spin_lock(&table_lock);
|
||||
hlist_for_each_entry_safe(entry, temp, &table_v4[i], hash) {
|
||||
if (unlikely(!work) ||
|
||||
now - entry->last_time_ns > NSEC_PER_SEC)
|
||||
entry_uninit(entry);
|
||||
}
|
||||
#if IS_ENABLED(CONFIG_IPV6)
|
||||
hlist_for_each_entry_safe(entry, temp, &table_v6[i], hash) {
|
||||
if (unlikely(!work) ||
|
||||
now - entry->last_time_ns > NSEC_PER_SEC)
|
||||
entry_uninit(entry);
|
||||
}
|
||||
#endif
|
||||
spin_unlock(&table_lock);
|
||||
if (likely(work))
|
||||
cond_resched();
|
||||
}
|
||||
if (likely(work))
|
||||
queue_delayed_work(system_power_efficient_wq, &gc_work, HZ);
|
||||
}
|
||||
|
||||
bool wg_ratelimiter_allow(struct sk_buff *skb, struct net *net)
|
||||
{
|
||||
/* We only take the bottom half of the net pointer, so that we can hash
|
||||
* 3 words in the end. This way, siphash's len param fits into the final
|
||||
* u32, and we don't incur an extra round.
|
||||
*/
|
||||
const u32 net_word = (unsigned long)net;
|
||||
struct ratelimiter_entry *entry;
|
||||
struct hlist_head *bucket;
|
||||
u64 ip;
|
||||
|
||||
if (skb->protocol == htons(ETH_P_IP)) {
|
||||
ip = (u64 __force)ip_hdr(skb)->saddr;
|
||||
bucket = &table_v4[hsiphash_2u32(net_word, ip, &key) &
|
||||
(table_size - 1)];
|
||||
}
|
||||
#if IS_ENABLED(CONFIG_IPV6)
|
||||
else if (skb->protocol == htons(ETH_P_IPV6)) {
|
||||
/* Only use 64 bits, so as to ratelimit the whole /64. */
|
||||
memcpy(&ip, &ipv6_hdr(skb)->saddr, sizeof(ip));
|
||||
bucket = &table_v6[hsiphash_3u32(net_word, ip >> 32, ip, &key) &
|
||||
(table_size - 1)];
|
||||
}
|
||||
#endif
|
||||
else
|
||||
return false;
|
||||
rcu_read_lock();
|
||||
hlist_for_each_entry_rcu(entry, bucket, hash) {
|
||||
if (entry->net == net && entry->ip == ip) {
|
||||
u64 now, tokens;
|
||||
bool ret;
|
||||
/* Quasi-inspired by nft_limit.c, but this is actually a
|
||||
* slightly different algorithm. Namely, we incorporate
|
||||
* the burst as part of the maximum tokens, rather than
|
||||
* as part of the rate.
|
||||
*/
|
||||
spin_lock(&entry->lock);
|
||||
now = ktime_get_coarse_boottime_ns();
|
||||
tokens = min_t(u64, TOKEN_MAX,
|
||||
entry->tokens + now -
|
||||
entry->last_time_ns);
|
||||
entry->last_time_ns = now;
|
||||
ret = tokens >= PACKET_COST;
|
||||
entry->tokens = ret ? tokens - PACKET_COST : tokens;
|
||||
spin_unlock(&entry->lock);
|
||||
rcu_read_unlock();
|
||||
return ret;
|
||||
}
|
||||
}
|
||||
rcu_read_unlock();
|
||||
|
||||
if (atomic_inc_return(&total_entries) > max_entries)
|
||||
goto err_oom;
|
||||
|
||||
entry = kmem_cache_alloc(entry_cache, GFP_KERNEL);
|
||||
if (unlikely(!entry))
|
||||
goto err_oom;
|
||||
|
||||
entry->net = net;
|
||||
entry->ip = ip;
|
||||
INIT_HLIST_NODE(&entry->hash);
|
||||
spin_lock_init(&entry->lock);
|
||||
entry->last_time_ns = ktime_get_coarse_boottime_ns();
|
||||
entry->tokens = TOKEN_MAX - PACKET_COST;
|
||||
spin_lock(&table_lock);
|
||||
hlist_add_head_rcu(&entry->hash, bucket);
|
||||
spin_unlock(&table_lock);
|
||||
return true;
|
||||
|
||||
err_oom:
|
||||
atomic_dec(&total_entries);
|
||||
return false;
|
||||
}
|
||||
|
||||
int wg_ratelimiter_init(void)
|
||||
{
|
||||
mutex_lock(&init_lock);
|
||||
if (++init_refcnt != 1)
|
||||
goto out;
|
||||
|
||||
entry_cache = KMEM_CACHE(ratelimiter_entry, 0);
|
||||
if (!entry_cache)
|
||||
goto err;
|
||||
|
||||
/* xt_hashlimit.c uses a slightly different algorithm for ratelimiting,
|
||||
* but what it shares in common is that it uses a massive hashtable. So,
|
||||
* we borrow their wisdom about good table sizes on different systems
|
||||
* dependent on RAM. This calculation here comes from there.
|
||||
*/
|
||||
table_size = (totalram_pages() > (1U << 30) / PAGE_SIZE) ? 8192 :
|
||||
max_t(unsigned long, 16, roundup_pow_of_two(
|
||||
(totalram_pages() << PAGE_SHIFT) /
|
||||
(1U << 14) / sizeof(struct hlist_head)));
|
||||
max_entries = table_size * 8;
|
||||
|
||||
table_v4 = kvcalloc(table_size, sizeof(*table_v4), GFP_KERNEL);
|
||||
if (unlikely(!table_v4))
|
||||
goto err_kmemcache;
|
||||
|
||||
#if IS_ENABLED(CONFIG_IPV6)
|
||||
table_v6 = kvcalloc(table_size, sizeof(*table_v6), GFP_KERNEL);
|
||||
if (unlikely(!table_v6)) {
|
||||
kvfree(table_v4);
|
||||
goto err_kmemcache;
|
||||
}
|
||||
#endif
|
||||
|
||||
queue_delayed_work(system_power_efficient_wq, &gc_work, HZ);
|
||||
get_random_bytes(&key, sizeof(key));
|
||||
out:
|
||||
mutex_unlock(&init_lock);
|
||||
return 0;
|
||||
|
||||
err_kmemcache:
|
||||
kmem_cache_destroy(entry_cache);
|
||||
err:
|
||||
--init_refcnt;
|
||||
mutex_unlock(&init_lock);
|
||||
return -ENOMEM;
|
||||
}
|
||||
|
||||
void wg_ratelimiter_uninit(void)
|
||||
{
|
||||
mutex_lock(&init_lock);
|
||||
if (!init_refcnt || --init_refcnt)
|
||||
goto out;
|
||||
|
||||
cancel_delayed_work_sync(&gc_work);
|
||||
wg_ratelimiter_gc_entries(NULL);
|
||||
rcu_barrier();
|
||||
kvfree(table_v4);
|
||||
#if IS_ENABLED(CONFIG_IPV6)
|
||||
kvfree(table_v6);
|
||||
#endif
|
||||
kmem_cache_destroy(entry_cache);
|
||||
out:
|
||||
mutex_unlock(&init_lock);
|
||||
}
|
||||
|
||||
#include "selftest/ratelimiter.c"
|
||||
19
drivers/net/wireguard/ratelimiter.h
Normal file
19
drivers/net/wireguard/ratelimiter.h
Normal file
|
|
@ -0,0 +1,19 @@
|
|||
/* SPDX-License-Identifier: GPL-2.0 */
|
||||
/*
|
||||
* Copyright (C) 2015-2019 Jason A. Donenfeld <Jason@zx2c4.com>. All Rights Reserved.
|
||||
*/
|
||||
|
||||
#ifndef _WG_RATELIMITER_H
|
||||
#define _WG_RATELIMITER_H
|
||||
|
||||
#include <linux/skbuff.h>
|
||||
|
||||
int wg_ratelimiter_init(void);
|
||||
void wg_ratelimiter_uninit(void);
|
||||
bool wg_ratelimiter_allow(struct sk_buff *skb, struct net *net);
|
||||
|
||||
#ifdef DEBUG
|
||||
bool wg_ratelimiter_selftest(void);
|
||||
#endif
|
||||
|
||||
#endif /* _WG_RATELIMITER_H */
|
||||
593
drivers/net/wireguard/receive.c
Normal file
593
drivers/net/wireguard/receive.c
Normal file
|
|
@ -0,0 +1,593 @@
|
|||
// SPDX-License-Identifier: GPL-2.0
|
||||
/*
|
||||
* Copyright (C) 2015-2019 Jason A. Donenfeld <Jason@zx2c4.com>. All Rights Reserved.
|
||||
*/
|
||||
|
||||
#include "queueing.h"
|
||||
#include "device.h"
|
||||
#include "peer.h"
|
||||
#include "timers.h"
|
||||
#include "messages.h"
|
||||
#include "cookie.h"
|
||||
#include "socket.h"
|
||||
|
||||
#include <linux/ip.h>
|
||||
#include <linux/ipv6.h>
|
||||
#include <linux/udp.h>
|
||||
#include <net/ip_tunnels.h>
|
||||
|
||||
/* Must be called with bh disabled. */
|
||||
static void update_rx_stats(struct wg_peer *peer, size_t len)
|
||||
{
|
||||
struct pcpu_sw_netstats *tstats =
|
||||
get_cpu_ptr(peer->device->dev->tstats);
|
||||
|
||||
u64_stats_update_begin(&tstats->syncp);
|
||||
++tstats->rx_packets;
|
||||
tstats->rx_bytes += len;
|
||||
peer->rx_bytes += len;
|
||||
u64_stats_update_end(&tstats->syncp);
|
||||
put_cpu_ptr(tstats);
|
||||
}
|
||||
|
||||
#define SKB_TYPE_LE32(skb) (((struct message_header *)(skb)->data)->type)
|
||||
|
||||
static size_t validate_header_len(struct sk_buff *skb)
|
||||
{
|
||||
if (unlikely(skb->len < sizeof(struct message_header)))
|
||||
return 0;
|
||||
if (SKB_TYPE_LE32(skb) == cpu_to_le32(MESSAGE_DATA) &&
|
||||
skb->len >= MESSAGE_MINIMUM_LENGTH)
|
||||
return sizeof(struct message_data);
|
||||
if (SKB_TYPE_LE32(skb) == cpu_to_le32(MESSAGE_HANDSHAKE_INITIATION) &&
|
||||
skb->len == sizeof(struct message_handshake_initiation))
|
||||
return sizeof(struct message_handshake_initiation);
|
||||
if (SKB_TYPE_LE32(skb) == cpu_to_le32(MESSAGE_HANDSHAKE_RESPONSE) &&
|
||||
skb->len == sizeof(struct message_handshake_response))
|
||||
return sizeof(struct message_handshake_response);
|
||||
if (SKB_TYPE_LE32(skb) == cpu_to_le32(MESSAGE_HANDSHAKE_COOKIE) &&
|
||||
skb->len == sizeof(struct message_handshake_cookie))
|
||||
return sizeof(struct message_handshake_cookie);
|
||||
return 0;
|
||||
}
|
||||
|
||||
static int prepare_skb_header(struct sk_buff *skb, struct wg_device *wg)
|
||||
{
|
||||
size_t data_offset, data_len, header_len;
|
||||
struct udphdr *udp;
|
||||
|
||||
if (unlikely(!wg_check_packet_protocol(skb) ||
|
||||
skb_transport_header(skb) < skb->head ||
|
||||
(skb_transport_header(skb) + sizeof(struct udphdr)) >
|
||||
skb_tail_pointer(skb)))
|
||||
return -EINVAL; /* Bogus IP header */
|
||||
udp = udp_hdr(skb);
|
||||
data_offset = (u8 *)udp - skb->data;
|
||||
if (unlikely(data_offset > U16_MAX ||
|
||||
data_offset + sizeof(struct udphdr) > skb->len))
|
||||
/* Packet has offset at impossible location or isn't big enough
|
||||
* to have UDP fields.
|
||||
*/
|
||||
return -EINVAL;
|
||||
data_len = ntohs(udp->len);
|
||||
if (unlikely(data_len < sizeof(struct udphdr) ||
|
||||
data_len > skb->len - data_offset))
|
||||
/* UDP packet is reporting too small of a size or lying about
|
||||
* its size.
|
||||
*/
|
||||
return -EINVAL;
|
||||
data_len -= sizeof(struct udphdr);
|
||||
data_offset = (u8 *)udp + sizeof(struct udphdr) - skb->data;
|
||||
if (unlikely(!pskb_may_pull(skb,
|
||||
data_offset + sizeof(struct message_header)) ||
|
||||
pskb_trim(skb, data_len + data_offset) < 0))
|
||||
return -EINVAL;
|
||||
skb_pull(skb, data_offset);
|
||||
if (unlikely(skb->len != data_len))
|
||||
/* Final len does not agree with calculated len */
|
||||
return -EINVAL;
|
||||
header_len = validate_header_len(skb);
|
||||
if (unlikely(!header_len))
|
||||
return -EINVAL;
|
||||
__skb_push(skb, data_offset);
|
||||
if (unlikely(!pskb_may_pull(skb, data_offset + header_len)))
|
||||
return -EINVAL;
|
||||
__skb_pull(skb, data_offset);
|
||||
return 0;
|
||||
}
|
||||
|
||||
static void wg_receive_handshake_packet(struct wg_device *wg,
|
||||
struct sk_buff *skb)
|
||||
{
|
||||
enum cookie_mac_state mac_state;
|
||||
struct wg_peer *peer = NULL;
|
||||
/* This is global, so that our load calculation applies to the whole
|
||||
* system. We don't care about races with it at all.
|
||||
*/
|
||||
static u64 last_under_load;
|
||||
bool packet_needs_cookie;
|
||||
bool under_load;
|
||||
|
||||
if (SKB_TYPE_LE32(skb) == cpu_to_le32(MESSAGE_HANDSHAKE_COOKIE)) {
|
||||
net_dbg_skb_ratelimited("%s: Receiving cookie response from %pISpfsc\n",
|
||||
wg->dev->name, skb);
|
||||
wg_cookie_message_consume(
|
||||
(struct message_handshake_cookie *)skb->data, wg);
|
||||
return;
|
||||
}
|
||||
|
||||
under_load = atomic_read(&wg->handshake_queue_len) >=
|
||||
MAX_QUEUED_INCOMING_HANDSHAKES / 8;
|
||||
if (under_load) {
|
||||
last_under_load = ktime_get_coarse_boottime_ns();
|
||||
} else if (last_under_load) {
|
||||
under_load = !wg_birthdate_has_expired(last_under_load, 1);
|
||||
if (!under_load)
|
||||
last_under_load = 0;
|
||||
}
|
||||
mac_state = wg_cookie_validate_packet(&wg->cookie_checker, skb,
|
||||
under_load);
|
||||
if ((under_load && mac_state == VALID_MAC_WITH_COOKIE) ||
|
||||
(!under_load && mac_state == VALID_MAC_BUT_NO_COOKIE)) {
|
||||
packet_needs_cookie = false;
|
||||
} else if (under_load && mac_state == VALID_MAC_BUT_NO_COOKIE) {
|
||||
packet_needs_cookie = true;
|
||||
} else {
|
||||
net_dbg_skb_ratelimited("%s: Invalid MAC of handshake, dropping packet from %pISpfsc\n",
|
||||
wg->dev->name, skb);
|
||||
return;
|
||||
}
|
||||
|
||||
switch (SKB_TYPE_LE32(skb)) {
|
||||
case cpu_to_le32(MESSAGE_HANDSHAKE_INITIATION): {
|
||||
struct message_handshake_initiation *message =
|
||||
(struct message_handshake_initiation *)skb->data;
|
||||
|
||||
if (packet_needs_cookie) {
|
||||
wg_packet_send_handshake_cookie(wg, skb,
|
||||
message->sender_index);
|
||||
return;
|
||||
}
|
||||
peer = wg_noise_handshake_consume_initiation(message, wg);
|
||||
if (unlikely(!peer)) {
|
||||
net_dbg_skb_ratelimited("%s: Invalid handshake initiation from %pISpfsc\n",
|
||||
wg->dev->name, skb);
|
||||
return;
|
||||
}
|
||||
wg_socket_set_peer_endpoint_from_skb(peer, skb);
|
||||
net_dbg_ratelimited("%s: Receiving handshake initiation from peer %llu (%pISpfsc)\n",
|
||||
wg->dev->name, peer->internal_id,
|
||||
&peer->endpoint.addr);
|
||||
wg_packet_send_handshake_response(peer);
|
||||
break;
|
||||
}
|
||||
case cpu_to_le32(MESSAGE_HANDSHAKE_RESPONSE): {
|
||||
struct message_handshake_response *message =
|
||||
(struct message_handshake_response *)skb->data;
|
||||
|
||||
if (packet_needs_cookie) {
|
||||
wg_packet_send_handshake_cookie(wg, skb,
|
||||
message->sender_index);
|
||||
return;
|
||||
}
|
||||
peer = wg_noise_handshake_consume_response(message, wg);
|
||||
if (unlikely(!peer)) {
|
||||
net_dbg_skb_ratelimited("%s: Invalid handshake response from %pISpfsc\n",
|
||||
wg->dev->name, skb);
|
||||
return;
|
||||
}
|
||||
wg_socket_set_peer_endpoint_from_skb(peer, skb);
|
||||
net_dbg_ratelimited("%s: Receiving handshake response from peer %llu (%pISpfsc)\n",
|
||||
wg->dev->name, peer->internal_id,
|
||||
&peer->endpoint.addr);
|
||||
if (wg_noise_handshake_begin_session(&peer->handshake,
|
||||
&peer->keypairs)) {
|
||||
wg_timers_session_derived(peer);
|
||||
wg_timers_handshake_complete(peer);
|
||||
/* Calling this function will either send any existing
|
||||
* packets in the queue and not send a keepalive, which
|
||||
* is the best case, Or, if there's nothing in the
|
||||
* queue, it will send a keepalive, in order to give
|
||||
* immediate confirmation of the session.
|
||||
*/
|
||||
wg_packet_send_keepalive(peer);
|
||||
}
|
||||
break;
|
||||
}
|
||||
}
|
||||
|
||||
if (unlikely(!peer)) {
|
||||
WARN(1, "Somehow a wrong type of packet wound up in the handshake queue!\n");
|
||||
return;
|
||||
}
|
||||
|
||||
local_bh_disable();
|
||||
update_rx_stats(peer, skb->len);
|
||||
local_bh_enable();
|
||||
|
||||
wg_timers_any_authenticated_packet_received(peer);
|
||||
wg_timers_any_authenticated_packet_traversal(peer);
|
||||
wg_peer_put(peer);
|
||||
}
|
||||
|
||||
void wg_packet_handshake_receive_worker(struct work_struct *work)
|
||||
{
|
||||
struct crypt_queue *queue = container_of(work, struct multicore_worker, work)->ptr;
|
||||
struct wg_device *wg = container_of(queue, struct wg_device, handshake_queue);
|
||||
struct sk_buff *skb;
|
||||
|
||||
while ((skb = ptr_ring_consume_bh(&queue->ring)) != NULL) {
|
||||
wg_receive_handshake_packet(wg, skb);
|
||||
dev_kfree_skb(skb);
|
||||
atomic_dec(&wg->handshake_queue_len);
|
||||
cond_resched();
|
||||
}
|
||||
}
|
||||
|
||||
static void keep_key_fresh(struct wg_peer *peer)
|
||||
{
|
||||
struct noise_keypair *keypair;
|
||||
bool send;
|
||||
|
||||
if (peer->sent_lastminute_handshake)
|
||||
return;
|
||||
|
||||
rcu_read_lock_bh();
|
||||
keypair = rcu_dereference_bh(peer->keypairs.current_keypair);
|
||||
send = keypair && READ_ONCE(keypair->sending.is_valid) &&
|
||||
keypair->i_am_the_initiator &&
|
||||
wg_birthdate_has_expired(keypair->sending.birthdate,
|
||||
REJECT_AFTER_TIME - KEEPALIVE_TIMEOUT - REKEY_TIMEOUT);
|
||||
rcu_read_unlock_bh();
|
||||
|
||||
if (unlikely(send)) {
|
||||
peer->sent_lastminute_handshake = true;
|
||||
wg_packet_send_queued_handshake_initiation(peer, false);
|
||||
}
|
||||
}
|
||||
|
||||
static bool decrypt_packet(struct sk_buff *skb, struct noise_keypair *keypair)
|
||||
{
|
||||
struct scatterlist sg[MAX_SKB_FRAGS + 8];
|
||||
struct sk_buff *trailer;
|
||||
unsigned int offset;
|
||||
int num_frags;
|
||||
|
||||
if (unlikely(!keypair))
|
||||
return false;
|
||||
|
||||
if (unlikely(!READ_ONCE(keypair->receiving.is_valid) ||
|
||||
wg_birthdate_has_expired(keypair->receiving.birthdate, REJECT_AFTER_TIME) ||
|
||||
READ_ONCE(keypair->receiving_counter.counter) >= REJECT_AFTER_MESSAGES)) {
|
||||
WRITE_ONCE(keypair->receiving.is_valid, false);
|
||||
return false;
|
||||
}
|
||||
|
||||
PACKET_CB(skb)->nonce =
|
||||
le64_to_cpu(((struct message_data *)skb->data)->counter);
|
||||
|
||||
/* We ensure that the network header is part of the packet before we
|
||||
* call skb_cow_data, so that there's no chance that data is removed
|
||||
* from the skb, so that later we can extract the original endpoint.
|
||||
*/
|
||||
offset = skb->data - skb_network_header(skb);
|
||||
skb_push(skb, offset);
|
||||
num_frags = skb_cow_data(skb, 0, &trailer);
|
||||
offset += sizeof(struct message_data);
|
||||
skb_pull(skb, offset);
|
||||
if (unlikely(num_frags < 0 || num_frags > ARRAY_SIZE(sg)))
|
||||
return false;
|
||||
|
||||
sg_init_table(sg, num_frags);
|
||||
if (skb_to_sgvec(skb, sg, 0, skb->len) <= 0)
|
||||
return false;
|
||||
|
||||
if (!chacha20poly1305_decrypt_sg_inplace(sg, skb->len, NULL, 0,
|
||||
PACKET_CB(skb)->nonce,
|
||||
keypair->receiving.key))
|
||||
return false;
|
||||
|
||||
/* Another ugly situation of pushing and pulling the header so as to
|
||||
* keep endpoint information intact.
|
||||
*/
|
||||
skb_push(skb, offset);
|
||||
if (pskb_trim(skb, skb->len - noise_encrypted_len(0)))
|
||||
return false;
|
||||
skb_pull(skb, offset);
|
||||
|
||||
return true;
|
||||
}
|
||||
|
||||
/* This is RFC6479, a replay detection bitmap algorithm that avoids bitshifts */
|
||||
static bool counter_validate(struct noise_replay_counter *counter, u64 their_counter)
|
||||
{
|
||||
unsigned long index, index_current, top, i;
|
||||
bool ret = false;
|
||||
|
||||
spin_lock_bh(&counter->lock);
|
||||
|
||||
if (unlikely(counter->counter >= REJECT_AFTER_MESSAGES + 1 ||
|
||||
their_counter >= REJECT_AFTER_MESSAGES))
|
||||
goto out;
|
||||
|
||||
++their_counter;
|
||||
|
||||
if (unlikely((COUNTER_WINDOW_SIZE + their_counter) <
|
||||
counter->counter))
|
||||
goto out;
|
||||
|
||||
index = their_counter >> ilog2(BITS_PER_LONG);
|
||||
|
||||
if (likely(their_counter > counter->counter)) {
|
||||
index_current = counter->counter >> ilog2(BITS_PER_LONG);
|
||||
top = min_t(unsigned long, index - index_current,
|
||||
COUNTER_BITS_TOTAL / BITS_PER_LONG);
|
||||
for (i = 1; i <= top; ++i)
|
||||
counter->backtrack[(i + index_current) &
|
||||
((COUNTER_BITS_TOTAL / BITS_PER_LONG) - 1)] = 0;
|
||||
WRITE_ONCE(counter->counter, their_counter);
|
||||
}
|
||||
|
||||
index &= (COUNTER_BITS_TOTAL / BITS_PER_LONG) - 1;
|
||||
ret = !test_and_set_bit(their_counter & (BITS_PER_LONG - 1),
|
||||
&counter->backtrack[index]);
|
||||
|
||||
out:
|
||||
spin_unlock_bh(&counter->lock);
|
||||
return ret;
|
||||
}
|
||||
|
||||
#include "selftest/counter.c"
|
||||
|
||||
static void wg_packet_consume_data_done(struct wg_peer *peer,
|
||||
struct sk_buff *skb,
|
||||
struct endpoint *endpoint)
|
||||
{
|
||||
struct net_device *dev = peer->device->dev;
|
||||
unsigned int len, len_before_trim;
|
||||
struct wg_peer *routed_peer;
|
||||
|
||||
wg_socket_set_peer_endpoint(peer, endpoint);
|
||||
|
||||
if (unlikely(wg_noise_received_with_keypair(&peer->keypairs,
|
||||
PACKET_CB(skb)->keypair))) {
|
||||
wg_timers_handshake_complete(peer);
|
||||
wg_packet_send_staged_packets(peer);
|
||||
}
|
||||
|
||||
keep_key_fresh(peer);
|
||||
|
||||
wg_timers_any_authenticated_packet_received(peer);
|
||||
wg_timers_any_authenticated_packet_traversal(peer);
|
||||
|
||||
/* A packet with length 0 is a keepalive packet */
|
||||
if (unlikely(!skb->len)) {
|
||||
update_rx_stats(peer, message_data_len(0));
|
||||
net_dbg_ratelimited("%s: Receiving keepalive packet from peer %llu (%pISpfsc)\n",
|
||||
dev->name, peer->internal_id,
|
||||
&peer->endpoint.addr);
|
||||
goto packet_processed;
|
||||
}
|
||||
|
||||
wg_timers_data_received(peer);
|
||||
|
||||
if (unlikely(skb_network_header(skb) < skb->head))
|
||||
goto dishonest_packet_size;
|
||||
if (unlikely(!(pskb_network_may_pull(skb, sizeof(struct iphdr)) &&
|
||||
(ip_hdr(skb)->version == 4 ||
|
||||
(ip_hdr(skb)->version == 6 &&
|
||||
pskb_network_may_pull(skb, sizeof(struct ipv6hdr)))))))
|
||||
goto dishonest_packet_type;
|
||||
|
||||
skb->dev = dev;
|
||||
/* We've already verified the Poly1305 auth tag, which means this packet
|
||||
* was not modified in transit. We can therefore tell the networking
|
||||
* stack that all checksums of every layer of encapsulation have already
|
||||
* been checked "by the hardware" and therefore is unnecessary to check
|
||||
* again in software.
|
||||
*/
|
||||
skb->ip_summed = CHECKSUM_UNNECESSARY;
|
||||
skb->csum_level = ~0; /* All levels */
|
||||
skb->protocol = ip_tunnel_parse_protocol(skb);
|
||||
if (skb->protocol == htons(ETH_P_IP)) {
|
||||
len = ntohs(ip_hdr(skb)->tot_len);
|
||||
if (unlikely(len < sizeof(struct iphdr)))
|
||||
goto dishonest_packet_size;
|
||||
INET_ECN_decapsulate(skb, PACKET_CB(skb)->ds, ip_hdr(skb)->tos);
|
||||
} else if (skb->protocol == htons(ETH_P_IPV6)) {
|
||||
len = ntohs(ipv6_hdr(skb)->payload_len) +
|
||||
sizeof(struct ipv6hdr);
|
||||
INET_ECN_decapsulate(skb, PACKET_CB(skb)->ds, ipv6_get_dsfield(ipv6_hdr(skb)));
|
||||
} else {
|
||||
goto dishonest_packet_type;
|
||||
}
|
||||
|
||||
if (unlikely(len > skb->len))
|
||||
goto dishonest_packet_size;
|
||||
len_before_trim = skb->len;
|
||||
if (unlikely(pskb_trim(skb, len)))
|
||||
goto packet_processed;
|
||||
|
||||
routed_peer = wg_allowedips_lookup_src(&peer->device->peer_allowedips,
|
||||
skb);
|
||||
wg_peer_put(routed_peer); /* We don't need the extra reference. */
|
||||
|
||||
if (unlikely(routed_peer != peer))
|
||||
goto dishonest_packet_peer;
|
||||
|
||||
napi_gro_receive(&peer->napi, skb);
|
||||
update_rx_stats(peer, message_data_len(len_before_trim));
|
||||
return;
|
||||
|
||||
dishonest_packet_peer:
|
||||
net_dbg_skb_ratelimited("%s: Packet has unallowed src IP (%pISc) from peer %llu (%pISpfsc)\n",
|
||||
dev->name, skb, peer->internal_id,
|
||||
&peer->endpoint.addr);
|
||||
++dev->stats.rx_errors;
|
||||
++dev->stats.rx_frame_errors;
|
||||
goto packet_processed;
|
||||
dishonest_packet_type:
|
||||
net_dbg_ratelimited("%s: Packet is neither ipv4 nor ipv6 from peer %llu (%pISpfsc)\n",
|
||||
dev->name, peer->internal_id, &peer->endpoint.addr);
|
||||
++dev->stats.rx_errors;
|
||||
++dev->stats.rx_frame_errors;
|
||||
goto packet_processed;
|
||||
dishonest_packet_size:
|
||||
net_dbg_ratelimited("%s: Packet has incorrect size from peer %llu (%pISpfsc)\n",
|
||||
dev->name, peer->internal_id, &peer->endpoint.addr);
|
||||
++dev->stats.rx_errors;
|
||||
++dev->stats.rx_length_errors;
|
||||
goto packet_processed;
|
||||
packet_processed:
|
||||
dev_kfree_skb(skb);
|
||||
}
|
||||
|
||||
int wg_packet_rx_poll(struct napi_struct *napi, int budget)
|
||||
{
|
||||
struct wg_peer *peer = container_of(napi, struct wg_peer, napi);
|
||||
struct noise_keypair *keypair;
|
||||
struct endpoint endpoint;
|
||||
enum packet_state state;
|
||||
struct sk_buff *skb;
|
||||
int work_done = 0;
|
||||
bool free;
|
||||
|
||||
if (unlikely(budget <= 0))
|
||||
return 0;
|
||||
|
||||
while ((skb = wg_prev_queue_peek(&peer->rx_queue)) != NULL &&
|
||||
(state = atomic_read_acquire(&PACKET_CB(skb)->state)) !=
|
||||
PACKET_STATE_UNCRYPTED) {
|
||||
wg_prev_queue_drop_peeked(&peer->rx_queue);
|
||||
keypair = PACKET_CB(skb)->keypair;
|
||||
free = true;
|
||||
|
||||
if (unlikely(state != PACKET_STATE_CRYPTED))
|
||||
goto next;
|
||||
|
||||
if (unlikely(!counter_validate(&keypair->receiving_counter,
|
||||
PACKET_CB(skb)->nonce))) {
|
||||
net_dbg_ratelimited("%s: Packet has invalid nonce %llu (max %llu)\n",
|
||||
peer->device->dev->name,
|
||||
PACKET_CB(skb)->nonce,
|
||||
READ_ONCE(keypair->receiving_counter.counter));
|
||||
goto next;
|
||||
}
|
||||
|
||||
if (unlikely(wg_socket_endpoint_from_skb(&endpoint, skb)))
|
||||
goto next;
|
||||
|
||||
wg_reset_packet(skb, false);
|
||||
wg_packet_consume_data_done(peer, skb, &endpoint);
|
||||
free = false;
|
||||
|
||||
next:
|
||||
wg_noise_keypair_put(keypair, false);
|
||||
wg_peer_put(peer);
|
||||
if (unlikely(free))
|
||||
dev_kfree_skb(skb);
|
||||
|
||||
if (++work_done >= budget)
|
||||
break;
|
||||
}
|
||||
|
||||
if (work_done < budget)
|
||||
napi_complete_done(napi, work_done);
|
||||
|
||||
return work_done;
|
||||
}
|
||||
|
||||
void wg_packet_decrypt_worker(struct work_struct *work)
|
||||
{
|
||||
struct crypt_queue *queue = container_of(work, struct multicore_worker,
|
||||
work)->ptr;
|
||||
struct sk_buff *skb;
|
||||
|
||||
while ((skb = ptr_ring_consume_bh(&queue->ring)) != NULL) {
|
||||
enum packet_state state =
|
||||
likely(decrypt_packet(skb, PACKET_CB(skb)->keypair)) ?
|
||||
PACKET_STATE_CRYPTED : PACKET_STATE_DEAD;
|
||||
wg_queue_enqueue_per_peer_rx(skb, state);
|
||||
if (need_resched())
|
||||
cond_resched();
|
||||
}
|
||||
}
|
||||
|
||||
static void wg_packet_consume_data(struct wg_device *wg, struct sk_buff *skb)
|
||||
{
|
||||
__le32 idx = ((struct message_data *)skb->data)->key_idx;
|
||||
struct wg_peer *peer = NULL;
|
||||
int ret;
|
||||
|
||||
rcu_read_lock_bh();
|
||||
PACKET_CB(skb)->keypair =
|
||||
(struct noise_keypair *)wg_index_hashtable_lookup(
|
||||
wg->index_hashtable, INDEX_HASHTABLE_KEYPAIR, idx,
|
||||
&peer);
|
||||
if (unlikely(!wg_noise_keypair_get(PACKET_CB(skb)->keypair)))
|
||||
goto err_keypair;
|
||||
|
||||
if (unlikely(READ_ONCE(peer->is_dead)))
|
||||
goto err;
|
||||
|
||||
ret = wg_queue_enqueue_per_device_and_peer(&wg->decrypt_queue, &peer->rx_queue, skb,
|
||||
wg->packet_crypt_wq);
|
||||
if (unlikely(ret == -EPIPE))
|
||||
wg_queue_enqueue_per_peer_rx(skb, PACKET_STATE_DEAD);
|
||||
if (likely(!ret || ret == -EPIPE)) {
|
||||
rcu_read_unlock_bh();
|
||||
return;
|
||||
}
|
||||
err:
|
||||
wg_noise_keypair_put(PACKET_CB(skb)->keypair, false);
|
||||
err_keypair:
|
||||
rcu_read_unlock_bh();
|
||||
wg_peer_put(peer);
|
||||
dev_kfree_skb(skb);
|
||||
}
|
||||
|
||||
void wg_packet_receive(struct wg_device *wg, struct sk_buff *skb)
|
||||
{
|
||||
if (unlikely(prepare_skb_header(skb, wg) < 0))
|
||||
goto err;
|
||||
switch (SKB_TYPE_LE32(skb)) {
|
||||
case cpu_to_le32(MESSAGE_HANDSHAKE_INITIATION):
|
||||
case cpu_to_le32(MESSAGE_HANDSHAKE_RESPONSE):
|
||||
case cpu_to_le32(MESSAGE_HANDSHAKE_COOKIE): {
|
||||
int cpu, ret = -EBUSY;
|
||||
|
||||
if (unlikely(!rng_is_initialized()))
|
||||
goto drop;
|
||||
if (atomic_read(&wg->handshake_queue_len) > MAX_QUEUED_INCOMING_HANDSHAKES / 2) {
|
||||
if (spin_trylock_bh(&wg->handshake_queue.ring.producer_lock)) {
|
||||
ret = __ptr_ring_produce(&wg->handshake_queue.ring, skb);
|
||||
spin_unlock_bh(&wg->handshake_queue.ring.producer_lock);
|
||||
}
|
||||
} else
|
||||
ret = ptr_ring_produce_bh(&wg->handshake_queue.ring, skb);
|
||||
if (ret) {
|
||||
drop:
|
||||
net_dbg_skb_ratelimited("%s: Dropping handshake packet from %pISpfsc\n",
|
||||
wg->dev->name, skb);
|
||||
goto err;
|
||||
}
|
||||
atomic_inc(&wg->handshake_queue_len);
|
||||
cpu = wg_cpumask_next_online(&wg->handshake_queue.last_cpu);
|
||||
/* Queues up a call to packet_process_queued_handshake_packets(skb): */
|
||||
queue_work_on(cpu, wg->handshake_receive_wq,
|
||||
&per_cpu_ptr(wg->handshake_queue.worker, cpu)->work);
|
||||
break;
|
||||
}
|
||||
case cpu_to_le32(MESSAGE_DATA):
|
||||
PACKET_CB(skb)->ds = ip_tunnel_get_dsfield(ip_hdr(skb), skb);
|
||||
wg_packet_consume_data(wg, skb);
|
||||
break;
|
||||
default:
|
||||
WARN(1, "Non-exhaustive parsing of packet header lead to unknown packet type!\n");
|
||||
goto err;
|
||||
}
|
||||
return;
|
||||
|
||||
err:
|
||||
dev_kfree_skb(skb);
|
||||
}
|
||||
680
drivers/net/wireguard/selftest/allowedips.c
Normal file
680
drivers/net/wireguard/selftest/allowedips.c
Normal file
|
|
@ -0,0 +1,680 @@
|
|||
// SPDX-License-Identifier: GPL-2.0
|
||||
/*
|
||||
* Copyright (C) 2015-2019 Jason A. Donenfeld <Jason@zx2c4.com>. All Rights Reserved.
|
||||
*
|
||||
* This contains some basic static unit tests for the allowedips data structure.
|
||||
* It also has two additional modes that are disabled and meant to be used by
|
||||
* folks directly playing with this file. If you define the macro
|
||||
* DEBUG_PRINT_TRIE_GRAPHVIZ to be 1, then every time there's a full tree in
|
||||
* memory, it will be printed out as KERN_DEBUG in a format that can be passed
|
||||
* to graphviz (the dot command) to visualize it. If you define the macro
|
||||
* DEBUG_RANDOM_TRIE to be 1, then there will be an extremely costly set of
|
||||
* randomized tests done against a trivial implementation, which may take
|
||||
* upwards of a half-hour to complete. There's no set of users who should be
|
||||
* enabling these, and the only developers that should go anywhere near these
|
||||
* nobs are the ones who are reading this comment.
|
||||
*/
|
||||
|
||||
#ifdef DEBUG
|
||||
|
||||
#include <linux/siphash.h>
|
||||
|
||||
static __init void print_node(struct allowedips_node *node, u8 bits)
|
||||
{
|
||||
char *fmt_connection = KERN_DEBUG "\t\"%p/%d\" -> \"%p/%d\";\n";
|
||||
char *fmt_declaration = KERN_DEBUG "\t\"%p/%d\"[style=%s, color=\"#%06x\"];\n";
|
||||
u8 ip1[16], ip2[16], cidr1, cidr2;
|
||||
char *style = "dotted";
|
||||
u32 color = 0;
|
||||
|
||||
if (node == NULL)
|
||||
return;
|
||||
if (bits == 32) {
|
||||
fmt_connection = KERN_DEBUG "\t\"%pI4/%d\" -> \"%pI4/%d\";\n";
|
||||
fmt_declaration = KERN_DEBUG "\t\"%pI4/%d\"[style=%s, color=\"#%06x\"];\n";
|
||||
} else if (bits == 128) {
|
||||
fmt_connection = KERN_DEBUG "\t\"%pI6/%d\" -> \"%pI6/%d\";\n";
|
||||
fmt_declaration = KERN_DEBUG "\t\"%pI6/%d\"[style=%s, color=\"#%06x\"];\n";
|
||||
}
|
||||
if (node->peer) {
|
||||
hsiphash_key_t key = { { 0 } };
|
||||
|
||||
memcpy(&key, &node->peer, sizeof(node->peer));
|
||||
color = hsiphash_1u32(0xdeadbeef, &key) % 200 << 16 |
|
||||
hsiphash_1u32(0xbabecafe, &key) % 200 << 8 |
|
||||
hsiphash_1u32(0xabad1dea, &key) % 200;
|
||||
style = "bold";
|
||||
}
|
||||
wg_allowedips_read_node(node, ip1, &cidr1);
|
||||
printk(fmt_declaration, ip1, cidr1, style, color);
|
||||
if (node->bit[0]) {
|
||||
wg_allowedips_read_node(rcu_dereference_raw(node->bit[0]), ip2, &cidr2);
|
||||
printk(fmt_connection, ip1, cidr1, ip2, cidr2);
|
||||
}
|
||||
if (node->bit[1]) {
|
||||
wg_allowedips_read_node(rcu_dereference_raw(node->bit[1]), ip2, &cidr2);
|
||||
printk(fmt_connection, ip1, cidr1, ip2, cidr2);
|
||||
}
|
||||
if (node->bit[0])
|
||||
print_node(rcu_dereference_raw(node->bit[0]), bits);
|
||||
if (node->bit[1])
|
||||
print_node(rcu_dereference_raw(node->bit[1]), bits);
|
||||
}
|
||||
|
||||
static __init void print_tree(struct allowedips_node __rcu *top, u8 bits)
|
||||
{
|
||||
printk(KERN_DEBUG "digraph trie {\n");
|
||||
print_node(rcu_dereference_raw(top), bits);
|
||||
printk(KERN_DEBUG "}\n");
|
||||
}
|
||||
|
||||
enum {
|
||||
NUM_PEERS = 2000,
|
||||
NUM_RAND_ROUTES = 400,
|
||||
NUM_MUTATED_ROUTES = 100,
|
||||
NUM_QUERIES = NUM_RAND_ROUTES * NUM_MUTATED_ROUTES * 30
|
||||
};
|
||||
|
||||
struct horrible_allowedips {
|
||||
struct hlist_head head;
|
||||
};
|
||||
|
||||
struct horrible_allowedips_node {
|
||||
struct hlist_node table;
|
||||
union nf_inet_addr ip;
|
||||
union nf_inet_addr mask;
|
||||
u8 ip_version;
|
||||
void *value;
|
||||
};
|
||||
|
||||
static __init void horrible_allowedips_init(struct horrible_allowedips *table)
|
||||
{
|
||||
INIT_HLIST_HEAD(&table->head);
|
||||
}
|
||||
|
||||
static __init void horrible_allowedips_free(struct horrible_allowedips *table)
|
||||
{
|
||||
struct horrible_allowedips_node *node;
|
||||
struct hlist_node *h;
|
||||
|
||||
hlist_for_each_entry_safe(node, h, &table->head, table) {
|
||||
hlist_del(&node->table);
|
||||
kfree(node);
|
||||
}
|
||||
}
|
||||
|
||||
static __init inline union nf_inet_addr horrible_cidr_to_mask(u8 cidr)
|
||||
{
|
||||
union nf_inet_addr mask;
|
||||
|
||||
memset(&mask, 0, sizeof(mask));
|
||||
memset(&mask.all, 0xff, cidr / 8);
|
||||
if (cidr % 32)
|
||||
mask.all[cidr / 32] = (__force u32)htonl(
|
||||
(0xFFFFFFFFUL << (32 - (cidr % 32))) & 0xFFFFFFFFUL);
|
||||
return mask;
|
||||
}
|
||||
|
||||
static __init inline u8 horrible_mask_to_cidr(union nf_inet_addr subnet)
|
||||
{
|
||||
return hweight32(subnet.all[0]) + hweight32(subnet.all[1]) +
|
||||
hweight32(subnet.all[2]) + hweight32(subnet.all[3]);
|
||||
}
|
||||
|
||||
static __init inline void
|
||||
horrible_mask_self(struct horrible_allowedips_node *node)
|
||||
{
|
||||
if (node->ip_version == 4) {
|
||||
node->ip.ip &= node->mask.ip;
|
||||
} else if (node->ip_version == 6) {
|
||||
node->ip.ip6[0] &= node->mask.ip6[0];
|
||||
node->ip.ip6[1] &= node->mask.ip6[1];
|
||||
node->ip.ip6[2] &= node->mask.ip6[2];
|
||||
node->ip.ip6[3] &= node->mask.ip6[3];
|
||||
}
|
||||
}
|
||||
|
||||
static __init inline bool
|
||||
horrible_match_v4(const struct horrible_allowedips_node *node, struct in_addr *ip)
|
||||
{
|
||||
return (ip->s_addr & node->mask.ip) == node->ip.ip;
|
||||
}
|
||||
|
||||
static __init inline bool
|
||||
horrible_match_v6(const struct horrible_allowedips_node *node, struct in6_addr *ip)
|
||||
{
|
||||
return (ip->in6_u.u6_addr32[0] & node->mask.ip6[0]) == node->ip.ip6[0] &&
|
||||
(ip->in6_u.u6_addr32[1] & node->mask.ip6[1]) == node->ip.ip6[1] &&
|
||||
(ip->in6_u.u6_addr32[2] & node->mask.ip6[2]) == node->ip.ip6[2] &&
|
||||
(ip->in6_u.u6_addr32[3] & node->mask.ip6[3]) == node->ip.ip6[3];
|
||||
}
|
||||
|
||||
static __init void
|
||||
horrible_insert_ordered(struct horrible_allowedips *table, struct horrible_allowedips_node *node)
|
||||
{
|
||||
struct horrible_allowedips_node *other = NULL, *where = NULL;
|
||||
u8 my_cidr = horrible_mask_to_cidr(node->mask);
|
||||
|
||||
hlist_for_each_entry(other, &table->head, table) {
|
||||
if (other->ip_version == node->ip_version &&
|
||||
!memcmp(&other->mask, &node->mask, sizeof(union nf_inet_addr)) &&
|
||||
!memcmp(&other->ip, &node->ip, sizeof(union nf_inet_addr))) {
|
||||
other->value = node->value;
|
||||
kfree(node);
|
||||
return;
|
||||
}
|
||||
}
|
||||
hlist_for_each_entry(other, &table->head, table) {
|
||||
where = other;
|
||||
if (horrible_mask_to_cidr(other->mask) <= my_cidr)
|
||||
break;
|
||||
}
|
||||
if (!other && !where)
|
||||
hlist_add_head(&node->table, &table->head);
|
||||
else if (!other)
|
||||
hlist_add_behind(&node->table, &where->table);
|
||||
else
|
||||
hlist_add_before(&node->table, &where->table);
|
||||
}
|
||||
|
||||
static __init int
|
||||
horrible_allowedips_insert_v4(struct horrible_allowedips *table,
|
||||
struct in_addr *ip, u8 cidr, void *value)
|
||||
{
|
||||
struct horrible_allowedips_node *node = kzalloc(sizeof(*node), GFP_KERNEL);
|
||||
|
||||
if (unlikely(!node))
|
||||
return -ENOMEM;
|
||||
node->ip.in = *ip;
|
||||
node->mask = horrible_cidr_to_mask(cidr);
|
||||
node->ip_version = 4;
|
||||
node->value = value;
|
||||
horrible_mask_self(node);
|
||||
horrible_insert_ordered(table, node);
|
||||
return 0;
|
||||
}
|
||||
|
||||
static __init int
|
||||
horrible_allowedips_insert_v6(struct horrible_allowedips *table,
|
||||
struct in6_addr *ip, u8 cidr, void *value)
|
||||
{
|
||||
struct horrible_allowedips_node *node = kzalloc(sizeof(*node), GFP_KERNEL);
|
||||
|
||||
if (unlikely(!node))
|
||||
return -ENOMEM;
|
||||
node->ip.in6 = *ip;
|
||||
node->mask = horrible_cidr_to_mask(cidr);
|
||||
node->ip_version = 6;
|
||||
node->value = value;
|
||||
horrible_mask_self(node);
|
||||
horrible_insert_ordered(table, node);
|
||||
return 0;
|
||||
}
|
||||
|
||||
static __init void *
|
||||
horrible_allowedips_lookup_v4(struct horrible_allowedips *table, struct in_addr *ip)
|
||||
{
|
||||
struct horrible_allowedips_node *node;
|
||||
|
||||
hlist_for_each_entry(node, &table->head, table) {
|
||||
if (node->ip_version == 4 && horrible_match_v4(node, ip))
|
||||
return node->value;
|
||||
}
|
||||
return NULL;
|
||||
}
|
||||
|
||||
static __init void *
|
||||
horrible_allowedips_lookup_v6(struct horrible_allowedips *table, struct in6_addr *ip)
|
||||
{
|
||||
struct horrible_allowedips_node *node;
|
||||
|
||||
hlist_for_each_entry(node, &table->head, table) {
|
||||
if (node->ip_version == 6 && horrible_match_v6(node, ip))
|
||||
return node->value;
|
||||
}
|
||||
return NULL;
|
||||
}
|
||||
|
||||
|
||||
static __init void
|
||||
horrible_allowedips_remove_by_value(struct horrible_allowedips *table, void *value)
|
||||
{
|
||||
struct horrible_allowedips_node *node;
|
||||
struct hlist_node *h;
|
||||
|
||||
hlist_for_each_entry_safe(node, h, &table->head, table) {
|
||||
if (node->value != value)
|
||||
continue;
|
||||
hlist_del(&node->table);
|
||||
kfree(node);
|
||||
}
|
||||
|
||||
}
|
||||
|
||||
static __init bool randomized_test(void)
|
||||
{
|
||||
unsigned int i, j, k, mutate_amount, cidr;
|
||||
u8 ip[16], mutate_mask[16], mutated[16];
|
||||
struct wg_peer **peers, *peer;
|
||||
struct horrible_allowedips h;
|
||||
DEFINE_MUTEX(mutex);
|
||||
struct allowedips t;
|
||||
bool ret = false;
|
||||
|
||||
mutex_init(&mutex);
|
||||
|
||||
wg_allowedips_init(&t);
|
||||
horrible_allowedips_init(&h);
|
||||
|
||||
peers = kcalloc(NUM_PEERS, sizeof(*peers), GFP_KERNEL);
|
||||
if (unlikely(!peers)) {
|
||||
pr_err("allowedips random self-test malloc: FAIL\n");
|
||||
goto free;
|
||||
}
|
||||
for (i = 0; i < NUM_PEERS; ++i) {
|
||||
peers[i] = kzalloc(sizeof(*peers[i]), GFP_KERNEL);
|
||||
if (unlikely(!peers[i])) {
|
||||
pr_err("allowedips random self-test malloc: FAIL\n");
|
||||
goto free;
|
||||
}
|
||||
kref_init(&peers[i]->refcount);
|
||||
INIT_LIST_HEAD(&peers[i]->allowedips_list);
|
||||
}
|
||||
|
||||
mutex_lock(&mutex);
|
||||
|
||||
for (i = 0; i < NUM_RAND_ROUTES; ++i) {
|
||||
prandom_bytes(ip, 4);
|
||||
cidr = prandom_u32_max(32) + 1;
|
||||
peer = peers[prandom_u32_max(NUM_PEERS)];
|
||||
if (wg_allowedips_insert_v4(&t, (struct in_addr *)ip, cidr,
|
||||
peer, &mutex) < 0) {
|
||||
pr_err("allowedips random self-test malloc: FAIL\n");
|
||||
goto free_locked;
|
||||
}
|
||||
if (horrible_allowedips_insert_v4(&h, (struct in_addr *)ip,
|
||||
cidr, peer) < 0) {
|
||||
pr_err("allowedips random self-test malloc: FAIL\n");
|
||||
goto free_locked;
|
||||
}
|
||||
for (j = 0; j < NUM_MUTATED_ROUTES; ++j) {
|
||||
memcpy(mutated, ip, 4);
|
||||
prandom_bytes(mutate_mask, 4);
|
||||
mutate_amount = prandom_u32_max(32);
|
||||
for (k = 0; k < mutate_amount / 8; ++k)
|
||||
mutate_mask[k] = 0xff;
|
||||
mutate_mask[k] = 0xff
|
||||
<< ((8 - (mutate_amount % 8)) % 8);
|
||||
for (; k < 4; ++k)
|
||||
mutate_mask[k] = 0;
|
||||
for (k = 0; k < 4; ++k)
|
||||
mutated[k] = (mutated[k] & mutate_mask[k]) |
|
||||
(~mutate_mask[k] &
|
||||
prandom_u32_max(256));
|
||||
cidr = prandom_u32_max(32) + 1;
|
||||
peer = peers[prandom_u32_max(NUM_PEERS)];
|
||||
if (wg_allowedips_insert_v4(&t,
|
||||
(struct in_addr *)mutated,
|
||||
cidr, peer, &mutex) < 0) {
|
||||
pr_err("allowedips random self-test malloc: FAIL\n");
|
||||
goto free_locked;
|
||||
}
|
||||
if (horrible_allowedips_insert_v4(&h,
|
||||
(struct in_addr *)mutated, cidr, peer)) {
|
||||
pr_err("allowedips random self-test malloc: FAIL\n");
|
||||
goto free_locked;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
for (i = 0; i < NUM_RAND_ROUTES; ++i) {
|
||||
prandom_bytes(ip, 16);
|
||||
cidr = prandom_u32_max(128) + 1;
|
||||
peer = peers[prandom_u32_max(NUM_PEERS)];
|
||||
if (wg_allowedips_insert_v6(&t, (struct in6_addr *)ip, cidr,
|
||||
peer, &mutex) < 0) {
|
||||
pr_err("allowedips random self-test malloc: FAIL\n");
|
||||
goto free_locked;
|
||||
}
|
||||
if (horrible_allowedips_insert_v6(&h, (struct in6_addr *)ip,
|
||||
cidr, peer) < 0) {
|
||||
pr_err("allowedips random self-test malloc: FAIL\n");
|
||||
goto free_locked;
|
||||
}
|
||||
for (j = 0; j < NUM_MUTATED_ROUTES; ++j) {
|
||||
memcpy(mutated, ip, 16);
|
||||
prandom_bytes(mutate_mask, 16);
|
||||
mutate_amount = prandom_u32_max(128);
|
||||
for (k = 0; k < mutate_amount / 8; ++k)
|
||||
mutate_mask[k] = 0xff;
|
||||
mutate_mask[k] = 0xff
|
||||
<< ((8 - (mutate_amount % 8)) % 8);
|
||||
for (; k < 4; ++k)
|
||||
mutate_mask[k] = 0;
|
||||
for (k = 0; k < 4; ++k)
|
||||
mutated[k] = (mutated[k] & mutate_mask[k]) |
|
||||
(~mutate_mask[k] &
|
||||
prandom_u32_max(256));
|
||||
cidr = prandom_u32_max(128) + 1;
|
||||
peer = peers[prandom_u32_max(NUM_PEERS)];
|
||||
if (wg_allowedips_insert_v6(&t,
|
||||
(struct in6_addr *)mutated,
|
||||
cidr, peer, &mutex) < 0) {
|
||||
pr_err("allowedips random self-test malloc: FAIL\n");
|
||||
goto free_locked;
|
||||
}
|
||||
if (horrible_allowedips_insert_v6(
|
||||
&h, (struct in6_addr *)mutated, cidr,
|
||||
peer)) {
|
||||
pr_err("allowedips random self-test malloc: FAIL\n");
|
||||
goto free_locked;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
mutex_unlock(&mutex);
|
||||
|
||||
if (IS_ENABLED(DEBUG_PRINT_TRIE_GRAPHVIZ)) {
|
||||
print_tree(t.root4, 32);
|
||||
print_tree(t.root6, 128);
|
||||
}
|
||||
|
||||
for (j = 0;; ++j) {
|
||||
for (i = 0; i < NUM_QUERIES; ++i) {
|
||||
prandom_bytes(ip, 4);
|
||||
if (lookup(t.root4, 32, ip) != horrible_allowedips_lookup_v4(&h, (struct in_addr *)ip)) {
|
||||
horrible_allowedips_lookup_v4(&h, (struct in_addr *)ip);
|
||||
pr_err("allowedips random v4 self-test: FAIL\n");
|
||||
goto free;
|
||||
}
|
||||
prandom_bytes(ip, 16);
|
||||
if (lookup(t.root6, 128, ip) != horrible_allowedips_lookup_v6(&h, (struct in6_addr *)ip)) {
|
||||
pr_err("allowedips random v6 self-test: FAIL\n");
|
||||
goto free;
|
||||
}
|
||||
}
|
||||
if (j >= NUM_PEERS)
|
||||
break;
|
||||
mutex_lock(&mutex);
|
||||
wg_allowedips_remove_by_peer(&t, peers[j], &mutex);
|
||||
mutex_unlock(&mutex);
|
||||
horrible_allowedips_remove_by_value(&h, peers[j]);
|
||||
}
|
||||
|
||||
if (t.root4 || t.root6) {
|
||||
pr_err("allowedips random self-test removal: FAIL\n");
|
||||
goto free;
|
||||
}
|
||||
|
||||
ret = true;
|
||||
|
||||
free:
|
||||
mutex_lock(&mutex);
|
||||
free_locked:
|
||||
wg_allowedips_free(&t, &mutex);
|
||||
mutex_unlock(&mutex);
|
||||
horrible_allowedips_free(&h);
|
||||
if (peers) {
|
||||
for (i = 0; i < NUM_PEERS; ++i)
|
||||
kfree(peers[i]);
|
||||
}
|
||||
kfree(peers);
|
||||
return ret;
|
||||
}
|
||||
|
||||
static __init inline struct in_addr *ip4(u8 a, u8 b, u8 c, u8 d)
|
||||
{
|
||||
static struct in_addr ip;
|
||||
u8 *split = (u8 *)&ip;
|
||||
|
||||
split[0] = a;
|
||||
split[1] = b;
|
||||
split[2] = c;
|
||||
split[3] = d;
|
||||
return &ip;
|
||||
}
|
||||
|
||||
static __init inline struct in6_addr *ip6(u32 a, u32 b, u32 c, u32 d)
|
||||
{
|
||||
static struct in6_addr ip;
|
||||
__be32 *split = (__be32 *)&ip;
|
||||
|
||||
split[0] = cpu_to_be32(a);
|
||||
split[1] = cpu_to_be32(b);
|
||||
split[2] = cpu_to_be32(c);
|
||||
split[3] = cpu_to_be32(d);
|
||||
return &ip;
|
||||
}
|
||||
|
||||
static __init struct wg_peer *init_peer(void)
|
||||
{
|
||||
struct wg_peer *peer = kzalloc(sizeof(*peer), GFP_KERNEL);
|
||||
|
||||
if (!peer)
|
||||
return NULL;
|
||||
kref_init(&peer->refcount);
|
||||
INIT_LIST_HEAD(&peer->allowedips_list);
|
||||
return peer;
|
||||
}
|
||||
|
||||
#define insert(version, mem, ipa, ipb, ipc, ipd, cidr) \
|
||||
wg_allowedips_insert_v##version(&t, ip##version(ipa, ipb, ipc, ipd), \
|
||||
cidr, mem, &mutex)
|
||||
|
||||
#define maybe_fail() do { \
|
||||
++i; \
|
||||
if (!_s) { \
|
||||
pr_info("allowedips self-test %zu: FAIL\n", i); \
|
||||
success = false; \
|
||||
} \
|
||||
} while (0)
|
||||
|
||||
#define test(version, mem, ipa, ipb, ipc, ipd) do { \
|
||||
bool _s = lookup(t.root##version, (version) == 4 ? 32 : 128, \
|
||||
ip##version(ipa, ipb, ipc, ipd)) == (mem); \
|
||||
maybe_fail(); \
|
||||
} while (0)
|
||||
|
||||
#define test_negative(version, mem, ipa, ipb, ipc, ipd) do { \
|
||||
bool _s = lookup(t.root##version, (version) == 4 ? 32 : 128, \
|
||||
ip##version(ipa, ipb, ipc, ipd)) != (mem); \
|
||||
maybe_fail(); \
|
||||
} while (0)
|
||||
|
||||
#define test_boolean(cond) do { \
|
||||
bool _s = (cond); \
|
||||
maybe_fail(); \
|
||||
} while (0)
|
||||
|
||||
bool __init wg_allowedips_selftest(void)
|
||||
{
|
||||
bool found_a = false, found_b = false, found_c = false, found_d = false,
|
||||
found_e = false, found_other = false;
|
||||
struct wg_peer *a = init_peer(), *b = init_peer(), *c = init_peer(),
|
||||
*d = init_peer(), *e = init_peer(), *f = init_peer(),
|
||||
*g = init_peer(), *h = init_peer();
|
||||
struct allowedips_node *iter_node;
|
||||
bool success = false;
|
||||
struct allowedips t;
|
||||
DEFINE_MUTEX(mutex);
|
||||
struct in6_addr ip;
|
||||
size_t i = 0, count = 0;
|
||||
__be64 part;
|
||||
|
||||
mutex_init(&mutex);
|
||||
mutex_lock(&mutex);
|
||||
wg_allowedips_init(&t);
|
||||
|
||||
if (!a || !b || !c || !d || !e || !f || !g || !h) {
|
||||
pr_err("allowedips self-test malloc: FAIL\n");
|
||||
goto free;
|
||||
}
|
||||
|
||||
insert(4, a, 192, 168, 4, 0, 24);
|
||||
insert(4, b, 192, 168, 4, 4, 32);
|
||||
insert(4, c, 192, 168, 0, 0, 16);
|
||||
insert(4, d, 192, 95, 5, 64, 27);
|
||||
/* replaces previous entry, and maskself is required */
|
||||
insert(4, c, 192, 95, 5, 65, 27);
|
||||
insert(6, d, 0x26075300, 0x60006b00, 0, 0xc05f0543, 128);
|
||||
insert(6, c, 0x26075300, 0x60006b00, 0, 0, 64);
|
||||
insert(4, e, 0, 0, 0, 0, 0);
|
||||
insert(6, e, 0, 0, 0, 0, 0);
|
||||
/* replaces previous entry */
|
||||
insert(6, f, 0, 0, 0, 0, 0);
|
||||
insert(6, g, 0x24046800, 0, 0, 0, 32);
|
||||
/* maskself is required */
|
||||
insert(6, h, 0x24046800, 0x40040800, 0xdeadbeef, 0xdeadbeef, 64);
|
||||
insert(6, a, 0x24046800, 0x40040800, 0xdeadbeef, 0xdeadbeef, 128);
|
||||
insert(6, c, 0x24446800, 0x40e40800, 0xdeaebeef, 0xdefbeef, 128);
|
||||
insert(6, b, 0x24446800, 0xf0e40800, 0xeeaebeef, 0, 98);
|
||||
insert(4, g, 64, 15, 112, 0, 20);
|
||||
/* maskself is required */
|
||||
insert(4, h, 64, 15, 123, 211, 25);
|
||||
insert(4, a, 10, 0, 0, 0, 25);
|
||||
insert(4, b, 10, 0, 0, 128, 25);
|
||||
insert(4, a, 10, 1, 0, 0, 30);
|
||||
insert(4, b, 10, 1, 0, 4, 30);
|
||||
insert(4, c, 10, 1, 0, 8, 29);
|
||||
insert(4, d, 10, 1, 0, 16, 29);
|
||||
|
||||
if (IS_ENABLED(DEBUG_PRINT_TRIE_GRAPHVIZ)) {
|
||||
print_tree(t.root4, 32);
|
||||
print_tree(t.root6, 128);
|
||||
}
|
||||
|
||||
success = true;
|
||||
|
||||
test(4, a, 192, 168, 4, 20);
|
||||
test(4, a, 192, 168, 4, 0);
|
||||
test(4, b, 192, 168, 4, 4);
|
||||
test(4, c, 192, 168, 200, 182);
|
||||
test(4, c, 192, 95, 5, 68);
|
||||
test(4, e, 192, 95, 5, 96);
|
||||
test(6, d, 0x26075300, 0x60006b00, 0, 0xc05f0543);
|
||||
test(6, c, 0x26075300, 0x60006b00, 0, 0xc02e01ee);
|
||||
test(6, f, 0x26075300, 0x60006b01, 0, 0);
|
||||
test(6, g, 0x24046800, 0x40040806, 0, 0x1006);
|
||||
test(6, g, 0x24046800, 0x40040806, 0x1234, 0x5678);
|
||||
test(6, f, 0x240467ff, 0x40040806, 0x1234, 0x5678);
|
||||
test(6, f, 0x24046801, 0x40040806, 0x1234, 0x5678);
|
||||
test(6, h, 0x24046800, 0x40040800, 0x1234, 0x5678);
|
||||
test(6, h, 0x24046800, 0x40040800, 0, 0);
|
||||
test(6, h, 0x24046800, 0x40040800, 0x10101010, 0x10101010);
|
||||
test(6, a, 0x24046800, 0x40040800, 0xdeadbeef, 0xdeadbeef);
|
||||
test(4, g, 64, 15, 116, 26);
|
||||
test(4, g, 64, 15, 127, 3);
|
||||
test(4, g, 64, 15, 123, 1);
|
||||
test(4, h, 64, 15, 123, 128);
|
||||
test(4, h, 64, 15, 123, 129);
|
||||
test(4, a, 10, 0, 0, 52);
|
||||
test(4, b, 10, 0, 0, 220);
|
||||
test(4, a, 10, 1, 0, 2);
|
||||
test(4, b, 10, 1, 0, 6);
|
||||
test(4, c, 10, 1, 0, 10);
|
||||
test(4, d, 10, 1, 0, 20);
|
||||
|
||||
insert(4, a, 1, 0, 0, 0, 32);
|
||||
insert(4, a, 64, 0, 0, 0, 32);
|
||||
insert(4, a, 128, 0, 0, 0, 32);
|
||||
insert(4, a, 192, 0, 0, 0, 32);
|
||||
insert(4, a, 255, 0, 0, 0, 32);
|
||||
wg_allowedips_remove_by_peer(&t, a, &mutex);
|
||||
test_negative(4, a, 1, 0, 0, 0);
|
||||
test_negative(4, a, 64, 0, 0, 0);
|
||||
test_negative(4, a, 128, 0, 0, 0);
|
||||
test_negative(4, a, 192, 0, 0, 0);
|
||||
test_negative(4, a, 255, 0, 0, 0);
|
||||
|
||||
wg_allowedips_free(&t, &mutex);
|
||||
wg_allowedips_init(&t);
|
||||
insert(4, a, 192, 168, 0, 0, 16);
|
||||
insert(4, a, 192, 168, 0, 0, 24);
|
||||
wg_allowedips_remove_by_peer(&t, a, &mutex);
|
||||
test_negative(4, a, 192, 168, 0, 1);
|
||||
|
||||
/* These will hit the WARN_ON(len >= MAX_ALLOWEDIPS_DEPTH) in free_node
|
||||
* if something goes wrong.
|
||||
*/
|
||||
for (i = 0; i < 64; ++i) {
|
||||
part = cpu_to_be64(~0LLU << i);
|
||||
memset(&ip, 0xff, 8);
|
||||
memcpy((u8 *)&ip + 8, &part, 8);
|
||||
wg_allowedips_insert_v6(&t, &ip, 128, a, &mutex);
|
||||
memcpy(&ip, &part, 8);
|
||||
memset((u8 *)&ip + 8, 0, 8);
|
||||
wg_allowedips_insert_v6(&t, &ip, 128, a, &mutex);
|
||||
}
|
||||
memset(&ip, 0, 16);
|
||||
wg_allowedips_insert_v6(&t, &ip, 128, a, &mutex);
|
||||
wg_allowedips_free(&t, &mutex);
|
||||
|
||||
wg_allowedips_init(&t);
|
||||
insert(4, a, 192, 95, 5, 93, 27);
|
||||
insert(6, a, 0x26075300, 0x60006b00, 0, 0xc05f0543, 128);
|
||||
insert(4, a, 10, 1, 0, 20, 29);
|
||||
insert(6, a, 0x26075300, 0x6d8a6bf8, 0xdab1f1df, 0xc05f1523, 83);
|
||||
insert(6, a, 0x26075300, 0x6d8a6bf8, 0xdab1f1df, 0xc05f1523, 21);
|
||||
list_for_each_entry(iter_node, &a->allowedips_list, peer_list) {
|
||||
u8 cidr, ip[16] __aligned(__alignof(u64));
|
||||
int family = wg_allowedips_read_node(iter_node, ip, &cidr);
|
||||
|
||||
count++;
|
||||
|
||||
if (cidr == 27 && family == AF_INET &&
|
||||
!memcmp(ip, ip4(192, 95, 5, 64), sizeof(struct in_addr)))
|
||||
found_a = true;
|
||||
else if (cidr == 128 && family == AF_INET6 &&
|
||||
!memcmp(ip, ip6(0x26075300, 0x60006b00, 0, 0xc05f0543),
|
||||
sizeof(struct in6_addr)))
|
||||
found_b = true;
|
||||
else if (cidr == 29 && family == AF_INET &&
|
||||
!memcmp(ip, ip4(10, 1, 0, 16), sizeof(struct in_addr)))
|
||||
found_c = true;
|
||||
else if (cidr == 83 && family == AF_INET6 &&
|
||||
!memcmp(ip, ip6(0x26075300, 0x6d8a6bf8, 0xdab1e000, 0),
|
||||
sizeof(struct in6_addr)))
|
||||
found_d = true;
|
||||
else if (cidr == 21 && family == AF_INET6 &&
|
||||
!memcmp(ip, ip6(0x26075000, 0, 0, 0),
|
||||
sizeof(struct in6_addr)))
|
||||
found_e = true;
|
||||
else
|
||||
found_other = true;
|
||||
}
|
||||
test_boolean(count == 5);
|
||||
test_boolean(found_a);
|
||||
test_boolean(found_b);
|
||||
test_boolean(found_c);
|
||||
test_boolean(found_d);
|
||||
test_boolean(found_e);
|
||||
test_boolean(!found_other);
|
||||
|
||||
if (IS_ENABLED(DEBUG_RANDOM_TRIE) && success)
|
||||
success = randomized_test();
|
||||
|
||||
if (success)
|
||||
pr_info("allowedips self-tests: pass\n");
|
||||
|
||||
free:
|
||||
wg_allowedips_free(&t, &mutex);
|
||||
kfree(a);
|
||||
kfree(b);
|
||||
kfree(c);
|
||||
kfree(d);
|
||||
kfree(e);
|
||||
kfree(f);
|
||||
kfree(g);
|
||||
kfree(h);
|
||||
mutex_unlock(&mutex);
|
||||
|
||||
return success;
|
||||
}
|
||||
|
||||
#undef test_negative
|
||||
#undef test
|
||||
#undef remove
|
||||
#undef insert
|
||||
#undef init_peer
|
||||
|
||||
#endif
|
||||
111
drivers/net/wireguard/selftest/counter.c
Normal file
111
drivers/net/wireguard/selftest/counter.c
Normal file
|
|
@ -0,0 +1,111 @@
|
|||
// SPDX-License-Identifier: GPL-2.0
|
||||
/*
|
||||
* Copyright (C) 2015-2019 Jason A. Donenfeld <Jason@zx2c4.com>. All Rights Reserved.
|
||||
*/
|
||||
|
||||
#ifdef DEBUG
|
||||
bool __init wg_packet_counter_selftest(void)
|
||||
{
|
||||
struct noise_replay_counter *counter;
|
||||
unsigned int test_num = 0, i;
|
||||
bool success = true;
|
||||
|
||||
counter = kmalloc(sizeof(*counter), GFP_KERNEL);
|
||||
if (unlikely(!counter)) {
|
||||
pr_err("nonce counter self-test malloc: FAIL\n");
|
||||
return false;
|
||||
}
|
||||
|
||||
#define T_INIT do { \
|
||||
memset(counter, 0, sizeof(*counter)); \
|
||||
spin_lock_init(&counter->lock); \
|
||||
} while (0)
|
||||
#define T_LIM (COUNTER_WINDOW_SIZE + 1)
|
||||
#define T(n, v) do { \
|
||||
++test_num; \
|
||||
if (counter_validate(counter, n) != (v)) { \
|
||||
pr_err("nonce counter self-test %u: FAIL\n", \
|
||||
test_num); \
|
||||
success = false; \
|
||||
} \
|
||||
} while (0)
|
||||
|
||||
T_INIT;
|
||||
/* 1 */ T(0, true);
|
||||
/* 2 */ T(1, true);
|
||||
/* 3 */ T(1, false);
|
||||
/* 4 */ T(9, true);
|
||||
/* 5 */ T(8, true);
|
||||
/* 6 */ T(7, true);
|
||||
/* 7 */ T(7, false);
|
||||
/* 8 */ T(T_LIM, true);
|
||||
/* 9 */ T(T_LIM - 1, true);
|
||||
/* 10 */ T(T_LIM - 1, false);
|
||||
/* 11 */ T(T_LIM - 2, true);
|
||||
/* 12 */ T(2, true);
|
||||
/* 13 */ T(2, false);
|
||||
/* 14 */ T(T_LIM + 16, true);
|
||||
/* 15 */ T(3, false);
|
||||
/* 16 */ T(T_LIM + 16, false);
|
||||
/* 17 */ T(T_LIM * 4, true);
|
||||
/* 18 */ T(T_LIM * 4 - (T_LIM - 1), true);
|
||||
/* 19 */ T(10, false);
|
||||
/* 20 */ T(T_LIM * 4 - T_LIM, false);
|
||||
/* 21 */ T(T_LIM * 4 - (T_LIM + 1), false);
|
||||
/* 22 */ T(T_LIM * 4 - (T_LIM - 2), true);
|
||||
/* 23 */ T(T_LIM * 4 + 1 - T_LIM, false);
|
||||
/* 24 */ T(0, false);
|
||||
/* 25 */ T(REJECT_AFTER_MESSAGES, false);
|
||||
/* 26 */ T(REJECT_AFTER_MESSAGES - 1, true);
|
||||
/* 27 */ T(REJECT_AFTER_MESSAGES, false);
|
||||
/* 28 */ T(REJECT_AFTER_MESSAGES - 1, false);
|
||||
/* 29 */ T(REJECT_AFTER_MESSAGES - 2, true);
|
||||
/* 30 */ T(REJECT_AFTER_MESSAGES + 1, false);
|
||||
/* 31 */ T(REJECT_AFTER_MESSAGES + 2, false);
|
||||
/* 32 */ T(REJECT_AFTER_MESSAGES - 2, false);
|
||||
/* 33 */ T(REJECT_AFTER_MESSAGES - 3, true);
|
||||
/* 34 */ T(0, false);
|
||||
|
||||
T_INIT;
|
||||
for (i = 1; i <= COUNTER_WINDOW_SIZE; ++i)
|
||||
T(i, true);
|
||||
T(0, true);
|
||||
T(0, false);
|
||||
|
||||
T_INIT;
|
||||
for (i = 2; i <= COUNTER_WINDOW_SIZE + 1; ++i)
|
||||
T(i, true);
|
||||
T(1, true);
|
||||
T(0, false);
|
||||
|
||||
T_INIT;
|
||||
for (i = COUNTER_WINDOW_SIZE + 1; i-- > 0;)
|
||||
T(i, true);
|
||||
|
||||
T_INIT;
|
||||
for (i = COUNTER_WINDOW_SIZE + 2; i-- > 1;)
|
||||
T(i, true);
|
||||
T(0, false);
|
||||
|
||||
T_INIT;
|
||||
for (i = COUNTER_WINDOW_SIZE + 1; i-- > 1;)
|
||||
T(i, true);
|
||||
T(COUNTER_WINDOW_SIZE + 1, true);
|
||||
T(0, false);
|
||||
|
||||
T_INIT;
|
||||
for (i = COUNTER_WINDOW_SIZE + 1; i-- > 1;)
|
||||
T(i, true);
|
||||
T(0, true);
|
||||
T(COUNTER_WINDOW_SIZE + 1, true);
|
||||
|
||||
#undef T
|
||||
#undef T_LIM
|
||||
#undef T_INIT
|
||||
|
||||
if (success)
|
||||
pr_info("nonce counter self-tests: pass\n");
|
||||
kfree(counter);
|
||||
return success;
|
||||
}
|
||||
#endif
|
||||
224
drivers/net/wireguard/selftest/ratelimiter.c
Normal file
224
drivers/net/wireguard/selftest/ratelimiter.c
Normal file
|
|
@ -0,0 +1,224 @@
|
|||
// SPDX-License-Identifier: GPL-2.0
|
||||
/*
|
||||
* Copyright (C) 2015-2019 Jason A. Donenfeld <Jason@zx2c4.com>. All Rights Reserved.
|
||||
*/
|
||||
|
||||
#ifdef DEBUG
|
||||
|
||||
#include <linux/jiffies.h>
|
||||
|
||||
static const struct {
|
||||
bool result;
|
||||
unsigned int msec_to_sleep_before;
|
||||
} expected_results[] __initconst = {
|
||||
[0 ... PACKETS_BURSTABLE - 1] = { true, 0 },
|
||||
[PACKETS_BURSTABLE] = { false, 0 },
|
||||
[PACKETS_BURSTABLE + 1] = { true, MSEC_PER_SEC / PACKETS_PER_SECOND },
|
||||
[PACKETS_BURSTABLE + 2] = { false, 0 },
|
||||
[PACKETS_BURSTABLE + 3] = { true, (MSEC_PER_SEC / PACKETS_PER_SECOND) * 2 },
|
||||
[PACKETS_BURSTABLE + 4] = { true, 0 },
|
||||
[PACKETS_BURSTABLE + 5] = { false, 0 }
|
||||
};
|
||||
|
||||
static __init unsigned int maximum_jiffies_at_index(int index)
|
||||
{
|
||||
unsigned int total_msecs = 2 * MSEC_PER_SEC / PACKETS_PER_SECOND / 3;
|
||||
int i;
|
||||
|
||||
for (i = 0; i <= index; ++i)
|
||||
total_msecs += expected_results[i].msec_to_sleep_before;
|
||||
return msecs_to_jiffies(total_msecs);
|
||||
}
|
||||
|
||||
static __init int timings_test(struct sk_buff *skb4, struct iphdr *hdr4,
|
||||
struct sk_buff *skb6, struct ipv6hdr *hdr6,
|
||||
int *test)
|
||||
{
|
||||
unsigned long loop_start_time;
|
||||
int i;
|
||||
|
||||
wg_ratelimiter_gc_entries(NULL);
|
||||
rcu_barrier();
|
||||
loop_start_time = jiffies;
|
||||
|
||||
for (i = 0; i < ARRAY_SIZE(expected_results); ++i) {
|
||||
if (expected_results[i].msec_to_sleep_before)
|
||||
msleep(expected_results[i].msec_to_sleep_before);
|
||||
|
||||
if (time_is_before_jiffies(loop_start_time +
|
||||
maximum_jiffies_at_index(i)))
|
||||
return -ETIMEDOUT;
|
||||
if (wg_ratelimiter_allow(skb4, &init_net) !=
|
||||
expected_results[i].result)
|
||||
return -EXFULL;
|
||||
++(*test);
|
||||
|
||||
hdr4->saddr = htonl(ntohl(hdr4->saddr) + i + 1);
|
||||
if (time_is_before_jiffies(loop_start_time +
|
||||
maximum_jiffies_at_index(i)))
|
||||
return -ETIMEDOUT;
|
||||
if (!wg_ratelimiter_allow(skb4, &init_net))
|
||||
return -EXFULL;
|
||||
++(*test);
|
||||
|
||||
hdr4->saddr = htonl(ntohl(hdr4->saddr) - i - 1);
|
||||
|
||||
#if IS_ENABLED(CONFIG_IPV6)
|
||||
hdr6->saddr.in6_u.u6_addr32[2] = htonl(i);
|
||||
hdr6->saddr.in6_u.u6_addr32[3] = htonl(i);
|
||||
if (time_is_before_jiffies(loop_start_time +
|
||||
maximum_jiffies_at_index(i)))
|
||||
return -ETIMEDOUT;
|
||||
if (wg_ratelimiter_allow(skb6, &init_net) !=
|
||||
expected_results[i].result)
|
||||
return -EXFULL;
|
||||
++(*test);
|
||||
|
||||
hdr6->saddr.in6_u.u6_addr32[0] =
|
||||
htonl(ntohl(hdr6->saddr.in6_u.u6_addr32[0]) + i + 1);
|
||||
if (time_is_before_jiffies(loop_start_time +
|
||||
maximum_jiffies_at_index(i)))
|
||||
return -ETIMEDOUT;
|
||||
if (!wg_ratelimiter_allow(skb6, &init_net))
|
||||
return -EXFULL;
|
||||
++(*test);
|
||||
|
||||
hdr6->saddr.in6_u.u6_addr32[0] =
|
||||
htonl(ntohl(hdr6->saddr.in6_u.u6_addr32[0]) - i - 1);
|
||||
|
||||
if (time_is_before_jiffies(loop_start_time +
|
||||
maximum_jiffies_at_index(i)))
|
||||
return -ETIMEDOUT;
|
||||
#endif
|
||||
}
|
||||
return 0;
|
||||
}
|
||||
|
||||
static __init int capacity_test(struct sk_buff *skb4, struct iphdr *hdr4,
|
||||
int *test)
|
||||
{
|
||||
int i;
|
||||
|
||||
wg_ratelimiter_gc_entries(NULL);
|
||||
rcu_barrier();
|
||||
|
||||
if (atomic_read(&total_entries))
|
||||
return -EXFULL;
|
||||
++(*test);
|
||||
|
||||
for (i = 0; i <= max_entries; ++i) {
|
||||
hdr4->saddr = htonl(i);
|
||||
if (wg_ratelimiter_allow(skb4, &init_net) != (i != max_entries))
|
||||
return -EXFULL;
|
||||
++(*test);
|
||||
}
|
||||
return 0;
|
||||
}
|
||||
|
||||
bool __init wg_ratelimiter_selftest(void)
|
||||
{
|
||||
enum { TRIALS_BEFORE_GIVING_UP = 5000 };
|
||||
bool success = false;
|
||||
int test = 0, trials;
|
||||
struct sk_buff *skb4, *skb6 = NULL;
|
||||
struct iphdr *hdr4;
|
||||
struct ipv6hdr *hdr6 = NULL;
|
||||
|
||||
if (IS_ENABLED(CONFIG_KASAN) || IS_ENABLED(CONFIG_UBSAN))
|
||||
return true;
|
||||
|
||||
BUILD_BUG_ON(MSEC_PER_SEC % PACKETS_PER_SECOND != 0);
|
||||
|
||||
if (wg_ratelimiter_init())
|
||||
goto out;
|
||||
++test;
|
||||
if (wg_ratelimiter_init()) {
|
||||
wg_ratelimiter_uninit();
|
||||
goto out;
|
||||
}
|
||||
++test;
|
||||
if (wg_ratelimiter_init()) {
|
||||
wg_ratelimiter_uninit();
|
||||
wg_ratelimiter_uninit();
|
||||
goto out;
|
||||
}
|
||||
++test;
|
||||
|
||||
skb4 = alloc_skb(sizeof(struct iphdr), GFP_KERNEL);
|
||||
if (unlikely(!skb4))
|
||||
goto err_nofree;
|
||||
skb4->protocol = htons(ETH_P_IP);
|
||||
hdr4 = (struct iphdr *)skb_put(skb4, sizeof(*hdr4));
|
||||
hdr4->saddr = htonl(8182);
|
||||
skb_reset_network_header(skb4);
|
||||
++test;
|
||||
|
||||
#if IS_ENABLED(CONFIG_IPV6)
|
||||
skb6 = alloc_skb(sizeof(struct ipv6hdr), GFP_KERNEL);
|
||||
if (unlikely(!skb6)) {
|
||||
kfree_skb(skb4);
|
||||
goto err_nofree;
|
||||
}
|
||||
skb6->protocol = htons(ETH_P_IPV6);
|
||||
hdr6 = (struct ipv6hdr *)skb_put(skb6, sizeof(*hdr6));
|
||||
hdr6->saddr.in6_u.u6_addr32[0] = htonl(1212);
|
||||
hdr6->saddr.in6_u.u6_addr32[1] = htonl(289188);
|
||||
skb_reset_network_header(skb6);
|
||||
++test;
|
||||
#endif
|
||||
|
||||
for (trials = TRIALS_BEFORE_GIVING_UP; IS_ENABLED(DEBUG_RATELIMITER_TIMINGS);) {
|
||||
int test_count = 0, ret;
|
||||
|
||||
ret = timings_test(skb4, hdr4, skb6, hdr6, &test_count);
|
||||
if (ret == -ETIMEDOUT) {
|
||||
if (!trials--) {
|
||||
test += test_count;
|
||||
goto err;
|
||||
}
|
||||
continue;
|
||||
} else if (ret < 0) {
|
||||
test += test_count;
|
||||
goto err;
|
||||
} else {
|
||||
test += test_count;
|
||||
break;
|
||||
}
|
||||
}
|
||||
|
||||
for (trials = TRIALS_BEFORE_GIVING_UP;;) {
|
||||
int test_count = 0;
|
||||
|
||||
if (capacity_test(skb4, hdr4, &test_count) < 0) {
|
||||
if (!trials--) {
|
||||
test += test_count;
|
||||
goto err;
|
||||
}
|
||||
continue;
|
||||
}
|
||||
test += test_count;
|
||||
break;
|
||||
}
|
||||
|
||||
success = true;
|
||||
|
||||
err:
|
||||
kfree_skb(skb4);
|
||||
#if IS_ENABLED(CONFIG_IPV6)
|
||||
kfree_skb(skb6);
|
||||
#endif
|
||||
err_nofree:
|
||||
wg_ratelimiter_uninit();
|
||||
wg_ratelimiter_uninit();
|
||||
wg_ratelimiter_uninit();
|
||||
/* Uninit one extra time to check underflow detection. */
|
||||
wg_ratelimiter_uninit();
|
||||
out:
|
||||
if (success)
|
||||
pr_info("ratelimiter self-tests: pass\n");
|
||||
else
|
||||
pr_err("ratelimiter self-test %d: FAIL\n", test);
|
||||
|
||||
return success;
|
||||
}
|
||||
#endif
|
||||
413
drivers/net/wireguard/send.c
Normal file
413
drivers/net/wireguard/send.c
Normal file
|
|
@ -0,0 +1,413 @@
|
|||
// SPDX-License-Identifier: GPL-2.0
|
||||
/*
|
||||
* Copyright (C) 2015-2019 Jason A. Donenfeld <Jason@zx2c4.com>. All Rights Reserved.
|
||||
*/
|
||||
|
||||
#include "queueing.h"
|
||||
#include "timers.h"
|
||||
#include "device.h"
|
||||
#include "peer.h"
|
||||
#include "socket.h"
|
||||
#include "messages.h"
|
||||
#include "cookie.h"
|
||||
|
||||
#include <linux/uio.h>
|
||||
#include <linux/inetdevice.h>
|
||||
#include <linux/socket.h>
|
||||
#include <net/ip_tunnels.h>
|
||||
#include <net/udp.h>
|
||||
#include <net/sock.h>
|
||||
|
||||
static void wg_packet_send_handshake_initiation(struct wg_peer *peer)
|
||||
{
|
||||
struct message_handshake_initiation packet;
|
||||
|
||||
if (!wg_birthdate_has_expired(atomic64_read(&peer->last_sent_handshake),
|
||||
REKEY_TIMEOUT))
|
||||
return; /* This function is rate limited. */
|
||||
|
||||
atomic64_set(&peer->last_sent_handshake, ktime_get_coarse_boottime_ns());
|
||||
net_dbg_ratelimited("%s: Sending handshake initiation to peer %llu (%pISpfsc)\n",
|
||||
peer->device->dev->name, peer->internal_id,
|
||||
&peer->endpoint.addr);
|
||||
|
||||
if (wg_noise_handshake_create_initiation(&packet, &peer->handshake)) {
|
||||
wg_cookie_add_mac_to_packet(&packet, sizeof(packet), peer);
|
||||
wg_timers_any_authenticated_packet_traversal(peer);
|
||||
wg_timers_any_authenticated_packet_sent(peer);
|
||||
atomic64_set(&peer->last_sent_handshake,
|
||||
ktime_get_coarse_boottime_ns());
|
||||
wg_socket_send_buffer_to_peer(peer, &packet, sizeof(packet),
|
||||
HANDSHAKE_DSCP);
|
||||
wg_timers_handshake_initiated(peer);
|
||||
}
|
||||
}
|
||||
|
||||
void wg_packet_handshake_send_worker(struct work_struct *work)
|
||||
{
|
||||
struct wg_peer *peer = container_of(work, struct wg_peer,
|
||||
transmit_handshake_work);
|
||||
|
||||
wg_packet_send_handshake_initiation(peer);
|
||||
wg_peer_put(peer);
|
||||
}
|
||||
|
||||
void wg_packet_send_queued_handshake_initiation(struct wg_peer *peer,
|
||||
bool is_retry)
|
||||
{
|
||||
if (!is_retry)
|
||||
peer->timer_handshake_attempts = 0;
|
||||
|
||||
rcu_read_lock_bh();
|
||||
/* We check last_sent_handshake here in addition to the actual function
|
||||
* we're queueing up, so that we don't queue things if not strictly
|
||||
* necessary:
|
||||
*/
|
||||
if (!wg_birthdate_has_expired(atomic64_read(&peer->last_sent_handshake),
|
||||
REKEY_TIMEOUT) ||
|
||||
unlikely(READ_ONCE(peer->is_dead)))
|
||||
goto out;
|
||||
|
||||
wg_peer_get(peer);
|
||||
/* Queues up calling packet_send_queued_handshakes(peer), where we do a
|
||||
* peer_put(peer) after:
|
||||
*/
|
||||
if (!queue_work(peer->device->handshake_send_wq,
|
||||
&peer->transmit_handshake_work))
|
||||
/* If the work was already queued, we want to drop the
|
||||
* extra reference:
|
||||
*/
|
||||
wg_peer_put(peer);
|
||||
out:
|
||||
rcu_read_unlock_bh();
|
||||
}
|
||||
|
||||
void wg_packet_send_handshake_response(struct wg_peer *peer)
|
||||
{
|
||||
struct message_handshake_response packet;
|
||||
|
||||
atomic64_set(&peer->last_sent_handshake, ktime_get_coarse_boottime_ns());
|
||||
net_dbg_ratelimited("%s: Sending handshake response to peer %llu (%pISpfsc)\n",
|
||||
peer->device->dev->name, peer->internal_id,
|
||||
&peer->endpoint.addr);
|
||||
|
||||
if (wg_noise_handshake_create_response(&packet, &peer->handshake)) {
|
||||
wg_cookie_add_mac_to_packet(&packet, sizeof(packet), peer);
|
||||
if (wg_noise_handshake_begin_session(&peer->handshake,
|
||||
&peer->keypairs)) {
|
||||
wg_timers_session_derived(peer);
|
||||
wg_timers_any_authenticated_packet_traversal(peer);
|
||||
wg_timers_any_authenticated_packet_sent(peer);
|
||||
atomic64_set(&peer->last_sent_handshake,
|
||||
ktime_get_coarse_boottime_ns());
|
||||
wg_socket_send_buffer_to_peer(peer, &packet,
|
||||
sizeof(packet),
|
||||
HANDSHAKE_DSCP);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
void wg_packet_send_handshake_cookie(struct wg_device *wg,
|
||||
struct sk_buff *initiating_skb,
|
||||
__le32 sender_index)
|
||||
{
|
||||
struct message_handshake_cookie packet;
|
||||
|
||||
net_dbg_skb_ratelimited("%s: Sending cookie response for denied handshake message for %pISpfsc\n",
|
||||
wg->dev->name, initiating_skb);
|
||||
wg_cookie_message_create(&packet, initiating_skb, sender_index,
|
||||
&wg->cookie_checker);
|
||||
wg_socket_send_buffer_as_reply_to_skb(wg, initiating_skb, &packet,
|
||||
sizeof(packet));
|
||||
}
|
||||
|
||||
static void keep_key_fresh(struct wg_peer *peer)
|
||||
{
|
||||
struct noise_keypair *keypair;
|
||||
bool send;
|
||||
|
||||
rcu_read_lock_bh();
|
||||
keypair = rcu_dereference_bh(peer->keypairs.current_keypair);
|
||||
send = keypair && READ_ONCE(keypair->sending.is_valid) &&
|
||||
(atomic64_read(&keypair->sending_counter) > REKEY_AFTER_MESSAGES ||
|
||||
(keypair->i_am_the_initiator &&
|
||||
wg_birthdate_has_expired(keypair->sending.birthdate, REKEY_AFTER_TIME)));
|
||||
rcu_read_unlock_bh();
|
||||
|
||||
if (unlikely(send))
|
||||
wg_packet_send_queued_handshake_initiation(peer, false);
|
||||
}
|
||||
|
||||
static unsigned int calculate_skb_padding(struct sk_buff *skb)
|
||||
{
|
||||
unsigned int padded_size, last_unit = skb->len;
|
||||
|
||||
if (unlikely(!PACKET_CB(skb)->mtu))
|
||||
return ALIGN(last_unit, MESSAGE_PADDING_MULTIPLE) - last_unit;
|
||||
|
||||
/* We do this modulo business with the MTU, just in case the networking
|
||||
* layer gives us a packet that's bigger than the MTU. In that case, we
|
||||
* wouldn't want the final subtraction to overflow in the case of the
|
||||
* padded_size being clamped. Fortunately, that's very rarely the case,
|
||||
* so we optimize for that not happening.
|
||||
*/
|
||||
if (unlikely(last_unit > PACKET_CB(skb)->mtu))
|
||||
last_unit %= PACKET_CB(skb)->mtu;
|
||||
|
||||
padded_size = min(PACKET_CB(skb)->mtu,
|
||||
ALIGN(last_unit, MESSAGE_PADDING_MULTIPLE));
|
||||
return padded_size - last_unit;
|
||||
}
|
||||
|
||||
static bool encrypt_packet(struct sk_buff *skb, struct noise_keypair *keypair)
|
||||
{
|
||||
unsigned int padding_len, plaintext_len, trailer_len;
|
||||
struct scatterlist sg[MAX_SKB_FRAGS + 8];
|
||||
struct message_data *header;
|
||||
struct sk_buff *trailer;
|
||||
int num_frags;
|
||||
|
||||
/* Force hash calculation before encryption so that flow analysis is
|
||||
* consistent over the inner packet.
|
||||
*/
|
||||
skb_get_hash(skb);
|
||||
|
||||
/* Calculate lengths. */
|
||||
padding_len = calculate_skb_padding(skb);
|
||||
trailer_len = padding_len + noise_encrypted_len(0);
|
||||
plaintext_len = skb->len + padding_len;
|
||||
|
||||
/* Expand data section to have room for padding and auth tag. */
|
||||
num_frags = skb_cow_data(skb, trailer_len, &trailer);
|
||||
if (unlikely(num_frags < 0 || num_frags > ARRAY_SIZE(sg)))
|
||||
return false;
|
||||
|
||||
/* Set the padding to zeros, and make sure it and the auth tag are part
|
||||
* of the skb.
|
||||
*/
|
||||
memset(skb_tail_pointer(trailer), 0, padding_len);
|
||||
|
||||
/* Expand head section to have room for our header and the network
|
||||
* stack's headers.
|
||||
*/
|
||||
if (unlikely(skb_cow_head(skb, DATA_PACKET_HEAD_ROOM) < 0))
|
||||
return false;
|
||||
|
||||
/* Finalize checksum calculation for the inner packet, if required. */
|
||||
if (unlikely(skb->ip_summed == CHECKSUM_PARTIAL &&
|
||||
skb_checksum_help(skb)))
|
||||
return false;
|
||||
|
||||
/* Only after checksumming can we safely add on the padding at the end
|
||||
* and the header.
|
||||
*/
|
||||
skb_set_inner_network_header(skb, 0);
|
||||
header = (struct message_data *)skb_push(skb, sizeof(*header));
|
||||
header->header.type = cpu_to_le32(MESSAGE_DATA);
|
||||
header->key_idx = keypair->remote_index;
|
||||
header->counter = cpu_to_le64(PACKET_CB(skb)->nonce);
|
||||
pskb_put(skb, trailer, trailer_len);
|
||||
|
||||
/* Now we can encrypt the scattergather segments */
|
||||
sg_init_table(sg, num_frags);
|
||||
if (skb_to_sgvec(skb, sg, sizeof(struct message_data),
|
||||
noise_encrypted_len(plaintext_len)) <= 0)
|
||||
return false;
|
||||
return chacha20poly1305_encrypt_sg_inplace(sg, plaintext_len, NULL, 0,
|
||||
PACKET_CB(skb)->nonce,
|
||||
keypair->sending.key);
|
||||
}
|
||||
|
||||
void wg_packet_send_keepalive(struct wg_peer *peer)
|
||||
{
|
||||
struct sk_buff *skb;
|
||||
|
||||
if (skb_queue_empty_lockless(&peer->staged_packet_queue)) {
|
||||
skb = alloc_skb(DATA_PACKET_HEAD_ROOM + MESSAGE_MINIMUM_LENGTH,
|
||||
GFP_ATOMIC);
|
||||
if (unlikely(!skb))
|
||||
return;
|
||||
skb_reserve(skb, DATA_PACKET_HEAD_ROOM);
|
||||
skb->dev = peer->device->dev;
|
||||
PACKET_CB(skb)->mtu = skb->dev->mtu;
|
||||
skb_queue_tail(&peer->staged_packet_queue, skb);
|
||||
net_dbg_ratelimited("%s: Sending keepalive packet to peer %llu (%pISpfsc)\n",
|
||||
peer->device->dev->name, peer->internal_id,
|
||||
&peer->endpoint.addr);
|
||||
}
|
||||
|
||||
wg_packet_send_staged_packets(peer);
|
||||
}
|
||||
|
||||
static void wg_packet_create_data_done(struct wg_peer *peer, struct sk_buff *first)
|
||||
{
|
||||
struct sk_buff *skb, *next;
|
||||
bool is_keepalive, data_sent = false;
|
||||
|
||||
wg_timers_any_authenticated_packet_traversal(peer);
|
||||
wg_timers_any_authenticated_packet_sent(peer);
|
||||
skb_list_walk_safe(first, skb, next) {
|
||||
is_keepalive = skb->len == message_data_len(0);
|
||||
if (likely(!wg_socket_send_skb_to_peer(peer, skb,
|
||||
PACKET_CB(skb)->ds) && !is_keepalive))
|
||||
data_sent = true;
|
||||
}
|
||||
|
||||
if (likely(data_sent))
|
||||
wg_timers_data_sent(peer);
|
||||
|
||||
keep_key_fresh(peer);
|
||||
}
|
||||
|
||||
void wg_packet_tx_worker(struct work_struct *work)
|
||||
{
|
||||
struct wg_peer *peer = container_of(work, struct wg_peer, transmit_packet_work);
|
||||
struct noise_keypair *keypair;
|
||||
enum packet_state state;
|
||||
struct sk_buff *first;
|
||||
|
||||
while ((first = wg_prev_queue_peek(&peer->tx_queue)) != NULL &&
|
||||
(state = atomic_read_acquire(&PACKET_CB(first)->state)) !=
|
||||
PACKET_STATE_UNCRYPTED) {
|
||||
wg_prev_queue_drop_peeked(&peer->tx_queue);
|
||||
keypair = PACKET_CB(first)->keypair;
|
||||
|
||||
if (likely(state == PACKET_STATE_CRYPTED))
|
||||
wg_packet_create_data_done(peer, first);
|
||||
else
|
||||
kfree_skb_list(first);
|
||||
|
||||
wg_noise_keypair_put(keypair, false);
|
||||
wg_peer_put(peer);
|
||||
if (need_resched())
|
||||
cond_resched();
|
||||
}
|
||||
}
|
||||
|
||||
void wg_packet_encrypt_worker(struct work_struct *work)
|
||||
{
|
||||
struct crypt_queue *queue = container_of(work, struct multicore_worker,
|
||||
work)->ptr;
|
||||
struct sk_buff *first, *skb, *next;
|
||||
|
||||
while ((first = ptr_ring_consume_bh(&queue->ring)) != NULL) {
|
||||
enum packet_state state = PACKET_STATE_CRYPTED;
|
||||
|
||||
skb_list_walk_safe(first, skb, next) {
|
||||
if (likely(encrypt_packet(skb,
|
||||
PACKET_CB(first)->keypair))) {
|
||||
wg_reset_packet(skb, true);
|
||||
} else {
|
||||
state = PACKET_STATE_DEAD;
|
||||
break;
|
||||
}
|
||||
}
|
||||
wg_queue_enqueue_per_peer_tx(first, state);
|
||||
if (need_resched())
|
||||
cond_resched();
|
||||
}
|
||||
}
|
||||
|
||||
static void wg_packet_create_data(struct wg_peer *peer, struct sk_buff *first)
|
||||
{
|
||||
struct wg_device *wg = peer->device;
|
||||
int ret = -EINVAL;
|
||||
|
||||
rcu_read_lock_bh();
|
||||
if (unlikely(READ_ONCE(peer->is_dead)))
|
||||
goto err;
|
||||
|
||||
ret = wg_queue_enqueue_per_device_and_peer(&wg->encrypt_queue, &peer->tx_queue, first,
|
||||
wg->packet_crypt_wq);
|
||||
if (unlikely(ret == -EPIPE))
|
||||
wg_queue_enqueue_per_peer_tx(first, PACKET_STATE_DEAD);
|
||||
err:
|
||||
rcu_read_unlock_bh();
|
||||
if (likely(!ret || ret == -EPIPE))
|
||||
return;
|
||||
wg_noise_keypair_put(PACKET_CB(first)->keypair, false);
|
||||
wg_peer_put(peer);
|
||||
kfree_skb_list(first);
|
||||
}
|
||||
|
||||
void wg_packet_purge_staged_packets(struct wg_peer *peer)
|
||||
{
|
||||
spin_lock_bh(&peer->staged_packet_queue.lock);
|
||||
peer->device->dev->stats.tx_dropped += peer->staged_packet_queue.qlen;
|
||||
__skb_queue_purge(&peer->staged_packet_queue);
|
||||
spin_unlock_bh(&peer->staged_packet_queue.lock);
|
||||
}
|
||||
|
||||
void wg_packet_send_staged_packets(struct wg_peer *peer)
|
||||
{
|
||||
struct noise_keypair *keypair;
|
||||
struct sk_buff_head packets;
|
||||
struct sk_buff *skb;
|
||||
|
||||
/* Steal the current queue into our local one. */
|
||||
__skb_queue_head_init(&packets);
|
||||
spin_lock_bh(&peer->staged_packet_queue.lock);
|
||||
skb_queue_splice_init(&peer->staged_packet_queue, &packets);
|
||||
spin_unlock_bh(&peer->staged_packet_queue.lock);
|
||||
if (unlikely(skb_queue_empty(&packets)))
|
||||
return;
|
||||
|
||||
/* First we make sure we have a valid reference to a valid key. */
|
||||
rcu_read_lock_bh();
|
||||
keypair = wg_noise_keypair_get(
|
||||
rcu_dereference_bh(peer->keypairs.current_keypair));
|
||||
rcu_read_unlock_bh();
|
||||
if (unlikely(!keypair))
|
||||
goto out_nokey;
|
||||
if (unlikely(!READ_ONCE(keypair->sending.is_valid)))
|
||||
goto out_nokey;
|
||||
if (unlikely(wg_birthdate_has_expired(keypair->sending.birthdate,
|
||||
REJECT_AFTER_TIME)))
|
||||
goto out_invalid;
|
||||
|
||||
/* After we know we have a somewhat valid key, we now try to assign
|
||||
* nonces to all of the packets in the queue. If we can't assign nonces
|
||||
* for all of them, we just consider it a failure and wait for the next
|
||||
* handshake.
|
||||
*/
|
||||
skb_queue_walk(&packets, skb) {
|
||||
/* 0 for no outer TOS: no leak. TODO: at some later point, we
|
||||
* might consider using flowi->tos as outer instead.
|
||||
*/
|
||||
PACKET_CB(skb)->ds = ip_tunnel_ecn_encap(0, ip_hdr(skb), skb);
|
||||
PACKET_CB(skb)->nonce =
|
||||
atomic64_inc_return(&keypair->sending_counter) - 1;
|
||||
if (unlikely(PACKET_CB(skb)->nonce >= REJECT_AFTER_MESSAGES))
|
||||
goto out_invalid;
|
||||
}
|
||||
|
||||
packets.prev->next = NULL;
|
||||
wg_peer_get(keypair->entry.peer);
|
||||
PACKET_CB(packets.next)->keypair = keypair;
|
||||
wg_packet_create_data(peer, packets.next);
|
||||
return;
|
||||
|
||||
out_invalid:
|
||||
WRITE_ONCE(keypair->sending.is_valid, false);
|
||||
out_nokey:
|
||||
wg_noise_keypair_put(keypair, false);
|
||||
|
||||
/* We orphan the packets if we're waiting on a handshake, so that they
|
||||
* don't block a socket's pool.
|
||||
*/
|
||||
skb_queue_walk(&packets, skb)
|
||||
skb_orphan(skb);
|
||||
/* Then we put them back on the top of the queue. We're not too
|
||||
* concerned about accidentally getting things a little out of order if
|
||||
* packets are being added really fast, because this queue is for before
|
||||
* packets can even be sent and it's small anyway.
|
||||
*/
|
||||
spin_lock_bh(&peer->staged_packet_queue.lock);
|
||||
skb_queue_splice(&packets, &peer->staged_packet_queue);
|
||||
spin_unlock_bh(&peer->staged_packet_queue.lock);
|
||||
|
||||
/* If we're exiting because there's something wrong with the key, it
|
||||
* means we should initiate a new handshake.
|
||||
*/
|
||||
wg_packet_send_queued_handshake_initiation(peer, false);
|
||||
}
|
||||
437
drivers/net/wireguard/socket.c
Normal file
437
drivers/net/wireguard/socket.c
Normal file
|
|
@ -0,0 +1,437 @@
|
|||
// SPDX-License-Identifier: GPL-2.0
|
||||
/*
|
||||
* Copyright (C) 2015-2019 Jason A. Donenfeld <Jason@zx2c4.com>. All Rights Reserved.
|
||||
*/
|
||||
|
||||
#include "device.h"
|
||||
#include "peer.h"
|
||||
#include "socket.h"
|
||||
#include "queueing.h"
|
||||
#include "messages.h"
|
||||
|
||||
#include <linux/ctype.h>
|
||||
#include <linux/net.h>
|
||||
#include <linux/if_vlan.h>
|
||||
#include <linux/if_ether.h>
|
||||
#include <linux/inetdevice.h>
|
||||
#include <net/udp_tunnel.h>
|
||||
#include <net/ipv6.h>
|
||||
|
||||
static int send4(struct wg_device *wg, struct sk_buff *skb,
|
||||
struct endpoint *endpoint, u8 ds, struct dst_cache *cache)
|
||||
{
|
||||
struct flowi4 fl = {
|
||||
.saddr = endpoint->src4.s_addr,
|
||||
.daddr = endpoint->addr4.sin_addr.s_addr,
|
||||
.fl4_dport = endpoint->addr4.sin_port,
|
||||
.flowi4_mark = wg->fwmark,
|
||||
.flowi4_proto = IPPROTO_UDP
|
||||
};
|
||||
struct rtable *rt = NULL;
|
||||
struct sock *sock;
|
||||
int ret = 0;
|
||||
|
||||
skb_mark_not_on_list(skb);
|
||||
skb->dev = wg->dev;
|
||||
skb->mark = wg->fwmark;
|
||||
|
||||
rcu_read_lock_bh();
|
||||
sock = rcu_dereference_bh(wg->sock4);
|
||||
|
||||
if (unlikely(!sock)) {
|
||||
ret = -ENONET;
|
||||
goto err;
|
||||
}
|
||||
|
||||
fl.fl4_sport = inet_sk(sock)->inet_sport;
|
||||
|
||||
if (cache)
|
||||
rt = dst_cache_get_ip4(cache, &fl.saddr);
|
||||
|
||||
if (!rt) {
|
||||
security_sk_classify_flow(sock, flowi4_to_flowi(&fl));
|
||||
if (unlikely(!inet_confirm_addr(sock_net(sock), NULL, 0,
|
||||
fl.saddr, RT_SCOPE_HOST))) {
|
||||
endpoint->src4.s_addr = 0;
|
||||
*(__force __be32 *)&endpoint->src_if4 = 0;
|
||||
fl.saddr = 0;
|
||||
if (cache)
|
||||
dst_cache_reset(cache);
|
||||
}
|
||||
rt = ip_route_output_flow(sock_net(sock), &fl, sock);
|
||||
if (unlikely(endpoint->src_if4 && ((IS_ERR(rt) &&
|
||||
PTR_ERR(rt) == -EINVAL) || (!IS_ERR(rt) &&
|
||||
rt->dst.dev->ifindex != endpoint->src_if4)))) {
|
||||
endpoint->src4.s_addr = 0;
|
||||
*(__force __be32 *)&endpoint->src_if4 = 0;
|
||||
fl.saddr = 0;
|
||||
if (cache)
|
||||
dst_cache_reset(cache);
|
||||
if (!IS_ERR(rt))
|
||||
ip_rt_put(rt);
|
||||
rt = ip_route_output_flow(sock_net(sock), &fl, sock);
|
||||
}
|
||||
if (unlikely(IS_ERR(rt))) {
|
||||
ret = PTR_ERR(rt);
|
||||
net_dbg_ratelimited("%s: No route to %pISpfsc, error %d\n",
|
||||
wg->dev->name, &endpoint->addr, ret);
|
||||
goto err;
|
||||
}
|
||||
if (cache)
|
||||
dst_cache_set_ip4(cache, &rt->dst, fl.saddr);
|
||||
}
|
||||
|
||||
skb->ignore_df = 1;
|
||||
udp_tunnel_xmit_skb(rt, sock, skb, fl.saddr, fl.daddr, ds,
|
||||
ip4_dst_hoplimit(&rt->dst), 0, fl.fl4_sport,
|
||||
fl.fl4_dport, false, false);
|
||||
goto out;
|
||||
|
||||
err:
|
||||
kfree_skb(skb);
|
||||
out:
|
||||
rcu_read_unlock_bh();
|
||||
return ret;
|
||||
}
|
||||
|
||||
static int send6(struct wg_device *wg, struct sk_buff *skb,
|
||||
struct endpoint *endpoint, u8 ds, struct dst_cache *cache)
|
||||
{
|
||||
#if IS_ENABLED(CONFIG_IPV6)
|
||||
struct flowi6 fl = {
|
||||
.saddr = endpoint->src6,
|
||||
.daddr = endpoint->addr6.sin6_addr,
|
||||
.fl6_dport = endpoint->addr6.sin6_port,
|
||||
.flowi6_mark = wg->fwmark,
|
||||
.flowi6_oif = endpoint->addr6.sin6_scope_id,
|
||||
.flowi6_proto = IPPROTO_UDP
|
||||
/* TODO: addr->sin6_flowinfo */
|
||||
};
|
||||
struct dst_entry *dst = NULL;
|
||||
struct sock *sock;
|
||||
int ret = 0;
|
||||
|
||||
skb_mark_not_on_list(skb);
|
||||
skb->dev = wg->dev;
|
||||
skb->mark = wg->fwmark;
|
||||
|
||||
rcu_read_lock_bh();
|
||||
sock = rcu_dereference_bh(wg->sock6);
|
||||
|
||||
if (unlikely(!sock)) {
|
||||
ret = -ENONET;
|
||||
goto err;
|
||||
}
|
||||
|
||||
fl.fl6_sport = inet_sk(sock)->inet_sport;
|
||||
|
||||
if (cache)
|
||||
dst = dst_cache_get_ip6(cache, &fl.saddr);
|
||||
|
||||
if (!dst) {
|
||||
security_sk_classify_flow(sock, flowi6_to_flowi(&fl));
|
||||
if (unlikely(!ipv6_addr_any(&fl.saddr) &&
|
||||
!ipv6_chk_addr(sock_net(sock), &fl.saddr, NULL, 0))) {
|
||||
endpoint->src6 = fl.saddr = in6addr_any;
|
||||
if (cache)
|
||||
dst_cache_reset(cache);
|
||||
}
|
||||
dst = ipv6_stub->ipv6_dst_lookup_flow(sock_net(sock), sock, &fl,
|
||||
NULL);
|
||||
if (unlikely(IS_ERR(dst))) {
|
||||
ret = PTR_ERR(dst);
|
||||
net_dbg_ratelimited("%s: No route to %pISpfsc, error %d\n",
|
||||
wg->dev->name, &endpoint->addr, ret);
|
||||
goto err;
|
||||
}
|
||||
if (cache)
|
||||
dst_cache_set_ip6(cache, dst, &fl.saddr);
|
||||
}
|
||||
|
||||
skb->ignore_df = 1;
|
||||
udp_tunnel6_xmit_skb(dst, sock, skb, skb->dev, &fl.saddr, &fl.daddr, ds,
|
||||
ip6_dst_hoplimit(dst), 0, fl.fl6_sport,
|
||||
fl.fl6_dport, false);
|
||||
goto out;
|
||||
|
||||
err:
|
||||
kfree_skb(skb);
|
||||
out:
|
||||
rcu_read_unlock_bh();
|
||||
return ret;
|
||||
#else
|
||||
kfree_skb(skb);
|
||||
return -EAFNOSUPPORT;
|
||||
#endif
|
||||
}
|
||||
|
||||
int wg_socket_send_skb_to_peer(struct wg_peer *peer, struct sk_buff *skb, u8 ds)
|
||||
{
|
||||
size_t skb_len = skb->len;
|
||||
int ret = -EAFNOSUPPORT;
|
||||
|
||||
read_lock_bh(&peer->endpoint_lock);
|
||||
if (peer->endpoint.addr.sa_family == AF_INET)
|
||||
ret = send4(peer->device, skb, &peer->endpoint, ds,
|
||||
&peer->endpoint_cache);
|
||||
else if (peer->endpoint.addr.sa_family == AF_INET6)
|
||||
ret = send6(peer->device, skb, &peer->endpoint, ds,
|
||||
&peer->endpoint_cache);
|
||||
else
|
||||
dev_kfree_skb(skb);
|
||||
if (likely(!ret))
|
||||
peer->tx_bytes += skb_len;
|
||||
read_unlock_bh(&peer->endpoint_lock);
|
||||
|
||||
return ret;
|
||||
}
|
||||
|
||||
int wg_socket_send_buffer_to_peer(struct wg_peer *peer, void *buffer,
|
||||
size_t len, u8 ds)
|
||||
{
|
||||
struct sk_buff *skb = alloc_skb(len + SKB_HEADER_LEN, GFP_ATOMIC);
|
||||
|
||||
if (unlikely(!skb))
|
||||
return -ENOMEM;
|
||||
|
||||
skb_reserve(skb, SKB_HEADER_LEN);
|
||||
skb_set_inner_network_header(skb, 0);
|
||||
skb_put_data(skb, buffer, len);
|
||||
return wg_socket_send_skb_to_peer(peer, skb, ds);
|
||||
}
|
||||
|
||||
int wg_socket_send_buffer_as_reply_to_skb(struct wg_device *wg,
|
||||
struct sk_buff *in_skb, void *buffer,
|
||||
size_t len)
|
||||
{
|
||||
int ret = 0;
|
||||
struct sk_buff *skb;
|
||||
struct endpoint endpoint;
|
||||
|
||||
if (unlikely(!in_skb))
|
||||
return -EINVAL;
|
||||
ret = wg_socket_endpoint_from_skb(&endpoint, in_skb);
|
||||
if (unlikely(ret < 0))
|
||||
return ret;
|
||||
|
||||
skb = alloc_skb(len + SKB_HEADER_LEN, GFP_ATOMIC);
|
||||
if (unlikely(!skb))
|
||||
return -ENOMEM;
|
||||
skb_reserve(skb, SKB_HEADER_LEN);
|
||||
skb_set_inner_network_header(skb, 0);
|
||||
skb_put_data(skb, buffer, len);
|
||||
|
||||
if (endpoint.addr.sa_family == AF_INET)
|
||||
ret = send4(wg, skb, &endpoint, 0, NULL);
|
||||
else if (endpoint.addr.sa_family == AF_INET6)
|
||||
ret = send6(wg, skb, &endpoint, 0, NULL);
|
||||
/* No other possibilities if the endpoint is valid, which it is,
|
||||
* as we checked above.
|
||||
*/
|
||||
|
||||
return ret;
|
||||
}
|
||||
|
||||
int wg_socket_endpoint_from_skb(struct endpoint *endpoint,
|
||||
const struct sk_buff *skb)
|
||||
{
|
||||
memset(endpoint, 0, sizeof(*endpoint));
|
||||
if (skb->protocol == htons(ETH_P_IP)) {
|
||||
endpoint->addr4.sin_family = AF_INET;
|
||||
endpoint->addr4.sin_port = udp_hdr(skb)->source;
|
||||
endpoint->addr4.sin_addr.s_addr = ip_hdr(skb)->saddr;
|
||||
endpoint->src4.s_addr = ip_hdr(skb)->daddr;
|
||||
endpoint->src_if4 = skb->skb_iif;
|
||||
} else if (IS_ENABLED(CONFIG_IPV6) && skb->protocol == htons(ETH_P_IPV6)) {
|
||||
endpoint->addr6.sin6_family = AF_INET6;
|
||||
endpoint->addr6.sin6_port = udp_hdr(skb)->source;
|
||||
endpoint->addr6.sin6_addr = ipv6_hdr(skb)->saddr;
|
||||
endpoint->addr6.sin6_scope_id = ipv6_iface_scope_id(
|
||||
&ipv6_hdr(skb)->saddr, skb->skb_iif);
|
||||
endpoint->src6 = ipv6_hdr(skb)->daddr;
|
||||
} else {
|
||||
return -EINVAL;
|
||||
}
|
||||
return 0;
|
||||
}
|
||||
|
||||
static bool endpoint_eq(const struct endpoint *a, const struct endpoint *b)
|
||||
{
|
||||
return (a->addr.sa_family == AF_INET && b->addr.sa_family == AF_INET &&
|
||||
a->addr4.sin_port == b->addr4.sin_port &&
|
||||
a->addr4.sin_addr.s_addr == b->addr4.sin_addr.s_addr &&
|
||||
a->src4.s_addr == b->src4.s_addr && a->src_if4 == b->src_if4) ||
|
||||
(a->addr.sa_family == AF_INET6 &&
|
||||
b->addr.sa_family == AF_INET6 &&
|
||||
a->addr6.sin6_port == b->addr6.sin6_port &&
|
||||
ipv6_addr_equal(&a->addr6.sin6_addr, &b->addr6.sin6_addr) &&
|
||||
a->addr6.sin6_scope_id == b->addr6.sin6_scope_id &&
|
||||
ipv6_addr_equal(&a->src6, &b->src6)) ||
|
||||
unlikely(!a->addr.sa_family && !b->addr.sa_family);
|
||||
}
|
||||
|
||||
void wg_socket_set_peer_endpoint(struct wg_peer *peer,
|
||||
const struct endpoint *endpoint)
|
||||
{
|
||||
/* First we check unlocked, in order to optimize, since it's pretty rare
|
||||
* that an endpoint will change. If we happen to be mid-write, and two
|
||||
* CPUs wind up writing the same thing or something slightly different,
|
||||
* it doesn't really matter much either.
|
||||
*/
|
||||
if (endpoint_eq(endpoint, &peer->endpoint))
|
||||
return;
|
||||
write_lock_bh(&peer->endpoint_lock);
|
||||
if (endpoint->addr.sa_family == AF_INET) {
|
||||
peer->endpoint.addr4 = endpoint->addr4;
|
||||
peer->endpoint.src4 = endpoint->src4;
|
||||
peer->endpoint.src_if4 = endpoint->src_if4;
|
||||
} else if (IS_ENABLED(CONFIG_IPV6) && endpoint->addr.sa_family == AF_INET6) {
|
||||
peer->endpoint.addr6 = endpoint->addr6;
|
||||
peer->endpoint.src6 = endpoint->src6;
|
||||
} else {
|
||||
goto out;
|
||||
}
|
||||
dst_cache_reset(&peer->endpoint_cache);
|
||||
out:
|
||||
write_unlock_bh(&peer->endpoint_lock);
|
||||
}
|
||||
|
||||
void wg_socket_set_peer_endpoint_from_skb(struct wg_peer *peer,
|
||||
const struct sk_buff *skb)
|
||||
{
|
||||
struct endpoint endpoint;
|
||||
|
||||
if (!wg_socket_endpoint_from_skb(&endpoint, skb))
|
||||
wg_socket_set_peer_endpoint(peer, &endpoint);
|
||||
}
|
||||
|
||||
void wg_socket_clear_peer_endpoint_src(struct wg_peer *peer)
|
||||
{
|
||||
write_lock_bh(&peer->endpoint_lock);
|
||||
memset(&peer->endpoint.src6, 0, sizeof(peer->endpoint.src6));
|
||||
dst_cache_reset_now(&peer->endpoint_cache);
|
||||
write_unlock_bh(&peer->endpoint_lock);
|
||||
}
|
||||
|
||||
static int wg_receive(struct sock *sk, struct sk_buff *skb)
|
||||
{
|
||||
struct wg_device *wg;
|
||||
|
||||
if (unlikely(!sk))
|
||||
goto err;
|
||||
wg = sk->sk_user_data;
|
||||
if (unlikely(!wg))
|
||||
goto err;
|
||||
skb_mark_not_on_list(skb);
|
||||
wg_packet_receive(wg, skb);
|
||||
return 0;
|
||||
|
||||
err:
|
||||
kfree_skb(skb);
|
||||
return 0;
|
||||
}
|
||||
|
||||
static void sock_free(struct sock *sock)
|
||||
{
|
||||
if (unlikely(!sock))
|
||||
return;
|
||||
sk_clear_memalloc(sock);
|
||||
udp_tunnel_sock_release(sock->sk_socket);
|
||||
}
|
||||
|
||||
static void set_sock_opts(struct socket *sock)
|
||||
{
|
||||
sock->sk->sk_allocation = GFP_ATOMIC;
|
||||
sock->sk->sk_sndbuf = INT_MAX;
|
||||
sk_set_memalloc(sock->sk);
|
||||
}
|
||||
|
||||
int wg_socket_init(struct wg_device *wg, u16 port)
|
||||
{
|
||||
struct net *net;
|
||||
int ret;
|
||||
struct udp_tunnel_sock_cfg cfg = {
|
||||
.sk_user_data = wg,
|
||||
.encap_type = 1,
|
||||
.encap_rcv = wg_receive
|
||||
};
|
||||
struct socket *new4 = NULL, *new6 = NULL;
|
||||
struct udp_port_cfg port4 = {
|
||||
.family = AF_INET,
|
||||
.local_ip.s_addr = htonl(INADDR_ANY),
|
||||
.local_udp_port = htons(port),
|
||||
.use_udp_checksums = true
|
||||
};
|
||||
#if IS_ENABLED(CONFIG_IPV6)
|
||||
int retries = 0;
|
||||
struct udp_port_cfg port6 = {
|
||||
.family = AF_INET6,
|
||||
.local_ip6 = IN6ADDR_ANY_INIT,
|
||||
.use_udp6_tx_checksums = true,
|
||||
.use_udp6_rx_checksums = true,
|
||||
.ipv6_v6only = true
|
||||
};
|
||||
#endif
|
||||
|
||||
rcu_read_lock();
|
||||
net = rcu_dereference(wg->creating_net);
|
||||
net = net ? maybe_get_net(net) : NULL;
|
||||
rcu_read_unlock();
|
||||
if (unlikely(!net))
|
||||
return -ENONET;
|
||||
|
||||
#if IS_ENABLED(CONFIG_IPV6)
|
||||
retry:
|
||||
#endif
|
||||
|
||||
ret = udp_sock_create(net, &port4, &new4);
|
||||
if (ret < 0) {
|
||||
pr_err("%s: Could not create IPv4 socket\n", wg->dev->name);
|
||||
goto out;
|
||||
}
|
||||
set_sock_opts(new4);
|
||||
setup_udp_tunnel_sock(net, new4, &cfg);
|
||||
|
||||
#if IS_ENABLED(CONFIG_IPV6)
|
||||
if (ipv6_mod_enabled()) {
|
||||
port6.local_udp_port = inet_sk(new4->sk)->inet_sport;
|
||||
ret = udp_sock_create(net, &port6, &new6);
|
||||
if (ret < 0) {
|
||||
udp_tunnel_sock_release(new4);
|
||||
if (ret == -EADDRINUSE && !port && retries++ < 100)
|
||||
goto retry;
|
||||
pr_err("%s: Could not create IPv6 socket\n",
|
||||
wg->dev->name);
|
||||
goto out;
|
||||
}
|
||||
set_sock_opts(new6);
|
||||
setup_udp_tunnel_sock(net, new6, &cfg);
|
||||
}
|
||||
#endif
|
||||
|
||||
wg_socket_reinit(wg, new4->sk, new6 ? new6->sk : NULL);
|
||||
ret = 0;
|
||||
out:
|
||||
put_net(net);
|
||||
return ret;
|
||||
}
|
||||
|
||||
void wg_socket_reinit(struct wg_device *wg, struct sock *new4,
|
||||
struct sock *new6)
|
||||
{
|
||||
struct sock *old4, *old6;
|
||||
|
||||
mutex_lock(&wg->socket_update_lock);
|
||||
old4 = rcu_dereference_protected(wg->sock4,
|
||||
lockdep_is_held(&wg->socket_update_lock));
|
||||
old6 = rcu_dereference_protected(wg->sock6,
|
||||
lockdep_is_held(&wg->socket_update_lock));
|
||||
rcu_assign_pointer(wg->sock4, new4);
|
||||
rcu_assign_pointer(wg->sock6, new6);
|
||||
if (new4)
|
||||
wg->incoming_port = ntohs(inet_sk(new4)->inet_sport);
|
||||
mutex_unlock(&wg->socket_update_lock);
|
||||
synchronize_net();
|
||||
sock_free(old4);
|
||||
sock_free(old6);
|
||||
}
|
||||
44
drivers/net/wireguard/socket.h
Normal file
44
drivers/net/wireguard/socket.h
Normal file
|
|
@ -0,0 +1,44 @@
|
|||
/* SPDX-License-Identifier: GPL-2.0 */
|
||||
/*
|
||||
* Copyright (C) 2015-2019 Jason A. Donenfeld <Jason@zx2c4.com>. All Rights Reserved.
|
||||
*/
|
||||
|
||||
#ifndef _WG_SOCKET_H
|
||||
#define _WG_SOCKET_H
|
||||
|
||||
#include <linux/netdevice.h>
|
||||
#include <linux/udp.h>
|
||||
#include <linux/if_vlan.h>
|
||||
#include <linux/if_ether.h>
|
||||
|
||||
int wg_socket_init(struct wg_device *wg, u16 port);
|
||||
void wg_socket_reinit(struct wg_device *wg, struct sock *new4,
|
||||
struct sock *new6);
|
||||
int wg_socket_send_buffer_to_peer(struct wg_peer *peer, void *data,
|
||||
size_t len, u8 ds);
|
||||
int wg_socket_send_skb_to_peer(struct wg_peer *peer, struct sk_buff *skb,
|
||||
u8 ds);
|
||||
int wg_socket_send_buffer_as_reply_to_skb(struct wg_device *wg,
|
||||
struct sk_buff *in_skb,
|
||||
void *out_buffer, size_t len);
|
||||
|
||||
int wg_socket_endpoint_from_skb(struct endpoint *endpoint,
|
||||
const struct sk_buff *skb);
|
||||
void wg_socket_set_peer_endpoint(struct wg_peer *peer,
|
||||
const struct endpoint *endpoint);
|
||||
void wg_socket_set_peer_endpoint_from_skb(struct wg_peer *peer,
|
||||
const struct sk_buff *skb);
|
||||
void wg_socket_clear_peer_endpoint_src(struct wg_peer *peer);
|
||||
|
||||
#if defined(CONFIG_DYNAMIC_DEBUG) || defined(DEBUG)
|
||||
#define net_dbg_skb_ratelimited(fmt, dev, skb, ...) do { \
|
||||
struct endpoint __endpoint; \
|
||||
wg_socket_endpoint_from_skb(&__endpoint, skb); \
|
||||
net_dbg_ratelimited(fmt, dev, &__endpoint.addr, \
|
||||
##__VA_ARGS__); \
|
||||
} while (0)
|
||||
#else
|
||||
#define net_dbg_skb_ratelimited(fmt, skb, ...)
|
||||
#endif
|
||||
|
||||
#endif /* _WG_SOCKET_H */
|
||||
243
drivers/net/wireguard/timers.c
Normal file
243
drivers/net/wireguard/timers.c
Normal file
|
|
@ -0,0 +1,243 @@
|
|||
// SPDX-License-Identifier: GPL-2.0
|
||||
/*
|
||||
* Copyright (C) 2015-2019 Jason A. Donenfeld <Jason@zx2c4.com>. All Rights Reserved.
|
||||
*/
|
||||
|
||||
#include "timers.h"
|
||||
#include "device.h"
|
||||
#include "peer.h"
|
||||
#include "queueing.h"
|
||||
#include "socket.h"
|
||||
|
||||
/*
|
||||
* - Timer for retransmitting the handshake if we don't hear back after
|
||||
* `REKEY_TIMEOUT + jitter` ms.
|
||||
*
|
||||
* - Timer for sending empty packet if we have received a packet but after have
|
||||
* not sent one for `KEEPALIVE_TIMEOUT` ms.
|
||||
*
|
||||
* - Timer for initiating new handshake if we have sent a packet but after have
|
||||
* not received one (even empty) for `(KEEPALIVE_TIMEOUT + REKEY_TIMEOUT) +
|
||||
* jitter` ms.
|
||||
*
|
||||
* - Timer for zeroing out all ephemeral keys after `(REJECT_AFTER_TIME * 3)` ms
|
||||
* if no new keys have been received.
|
||||
*
|
||||
* - Timer for, if enabled, sending an empty authenticated packet every user-
|
||||
* specified seconds.
|
||||
*/
|
||||
|
||||
static inline void mod_peer_timer(struct wg_peer *peer,
|
||||
struct timer_list *timer,
|
||||
unsigned long expires)
|
||||
{
|
||||
rcu_read_lock_bh();
|
||||
if (likely(netif_running(peer->device->dev) &&
|
||||
!READ_ONCE(peer->is_dead)))
|
||||
mod_timer(timer, expires);
|
||||
rcu_read_unlock_bh();
|
||||
}
|
||||
|
||||
static void wg_expired_retransmit_handshake(struct timer_list *timer)
|
||||
{
|
||||
struct wg_peer *peer = from_timer(peer, timer,
|
||||
timer_retransmit_handshake);
|
||||
|
||||
if (peer->timer_handshake_attempts > MAX_TIMER_HANDSHAKES) {
|
||||
pr_debug("%s: Handshake for peer %llu (%pISpfsc) did not complete after %d attempts, giving up\n",
|
||||
peer->device->dev->name, peer->internal_id,
|
||||
&peer->endpoint.addr, MAX_TIMER_HANDSHAKES + 2);
|
||||
|
||||
del_timer(&peer->timer_send_keepalive);
|
||||
/* We drop all packets without a keypair and don't try again,
|
||||
* if we try unsuccessfully for too long to make a handshake.
|
||||
*/
|
||||
wg_packet_purge_staged_packets(peer);
|
||||
|
||||
/* We set a timer for destroying any residue that might be left
|
||||
* of a partial exchange.
|
||||
*/
|
||||
if (!timer_pending(&peer->timer_zero_key_material))
|
||||
mod_peer_timer(peer, &peer->timer_zero_key_material,
|
||||
jiffies + REJECT_AFTER_TIME * 3 * HZ);
|
||||
} else {
|
||||
++peer->timer_handshake_attempts;
|
||||
pr_debug("%s: Handshake for peer %llu (%pISpfsc) did not complete after %d seconds, retrying (try %d)\n",
|
||||
peer->device->dev->name, peer->internal_id,
|
||||
&peer->endpoint.addr, REKEY_TIMEOUT,
|
||||
peer->timer_handshake_attempts + 1);
|
||||
|
||||
/* We clear the endpoint address src address, in case this is
|
||||
* the cause of trouble.
|
||||
*/
|
||||
wg_socket_clear_peer_endpoint_src(peer);
|
||||
|
||||
wg_packet_send_queued_handshake_initiation(peer, true);
|
||||
}
|
||||
}
|
||||
|
||||
static void wg_expired_send_keepalive(struct timer_list *timer)
|
||||
{
|
||||
struct wg_peer *peer = from_timer(peer, timer, timer_send_keepalive);
|
||||
|
||||
wg_packet_send_keepalive(peer);
|
||||
if (peer->timer_need_another_keepalive) {
|
||||
peer->timer_need_another_keepalive = false;
|
||||
mod_peer_timer(peer, &peer->timer_send_keepalive,
|
||||
jiffies + KEEPALIVE_TIMEOUT * HZ);
|
||||
}
|
||||
}
|
||||
|
||||
static void wg_expired_new_handshake(struct timer_list *timer)
|
||||
{
|
||||
struct wg_peer *peer = from_timer(peer, timer, timer_new_handshake);
|
||||
|
||||
pr_debug("%s: Retrying handshake with peer %llu (%pISpfsc) because we stopped hearing back after %d seconds\n",
|
||||
peer->device->dev->name, peer->internal_id,
|
||||
&peer->endpoint.addr, KEEPALIVE_TIMEOUT + REKEY_TIMEOUT);
|
||||
/* We clear the endpoint address src address, in case this is the cause
|
||||
* of trouble.
|
||||
*/
|
||||
wg_socket_clear_peer_endpoint_src(peer);
|
||||
wg_packet_send_queued_handshake_initiation(peer, false);
|
||||
}
|
||||
|
||||
static void wg_expired_zero_key_material(struct timer_list *timer)
|
||||
{
|
||||
struct wg_peer *peer = from_timer(peer, timer, timer_zero_key_material);
|
||||
|
||||
rcu_read_lock_bh();
|
||||
if (!READ_ONCE(peer->is_dead)) {
|
||||
wg_peer_get(peer);
|
||||
if (!queue_work(peer->device->handshake_send_wq,
|
||||
&peer->clear_peer_work))
|
||||
/* If the work was already on the queue, we want to drop
|
||||
* the extra reference.
|
||||
*/
|
||||
wg_peer_put(peer);
|
||||
}
|
||||
rcu_read_unlock_bh();
|
||||
}
|
||||
|
||||
static void wg_queued_expired_zero_key_material(struct work_struct *work)
|
||||
{
|
||||
struct wg_peer *peer = container_of(work, struct wg_peer,
|
||||
clear_peer_work);
|
||||
|
||||
pr_debug("%s: Zeroing out all keys for peer %llu (%pISpfsc), since we haven't received a new one in %d seconds\n",
|
||||
peer->device->dev->name, peer->internal_id,
|
||||
&peer->endpoint.addr, REJECT_AFTER_TIME * 3);
|
||||
wg_noise_handshake_clear(&peer->handshake);
|
||||
wg_noise_keypairs_clear(&peer->keypairs);
|
||||
wg_peer_put(peer);
|
||||
}
|
||||
|
||||
static void wg_expired_send_persistent_keepalive(struct timer_list *timer)
|
||||
{
|
||||
struct wg_peer *peer = from_timer(peer, timer,
|
||||
timer_persistent_keepalive);
|
||||
|
||||
if (likely(peer->persistent_keepalive_interval))
|
||||
wg_packet_send_keepalive(peer);
|
||||
}
|
||||
|
||||
/* Should be called after an authenticated data packet is sent. */
|
||||
void wg_timers_data_sent(struct wg_peer *peer)
|
||||
{
|
||||
if (!timer_pending(&peer->timer_new_handshake))
|
||||
mod_peer_timer(peer, &peer->timer_new_handshake,
|
||||
jiffies + (KEEPALIVE_TIMEOUT + REKEY_TIMEOUT) * HZ +
|
||||
prandom_u32_max(REKEY_TIMEOUT_JITTER_MAX_JIFFIES));
|
||||
}
|
||||
|
||||
/* Should be called after an authenticated data packet is received. */
|
||||
void wg_timers_data_received(struct wg_peer *peer)
|
||||
{
|
||||
if (likely(netif_running(peer->device->dev))) {
|
||||
if (!timer_pending(&peer->timer_send_keepalive))
|
||||
mod_peer_timer(peer, &peer->timer_send_keepalive,
|
||||
jiffies + KEEPALIVE_TIMEOUT * HZ);
|
||||
else
|
||||
peer->timer_need_another_keepalive = true;
|
||||
}
|
||||
}
|
||||
|
||||
/* Should be called after any type of authenticated packet is sent, whether
|
||||
* keepalive, data, or handshake.
|
||||
*/
|
||||
void wg_timers_any_authenticated_packet_sent(struct wg_peer *peer)
|
||||
{
|
||||
del_timer(&peer->timer_send_keepalive);
|
||||
}
|
||||
|
||||
/* Should be called after any type of authenticated packet is received, whether
|
||||
* keepalive, data, or handshake.
|
||||
*/
|
||||
void wg_timers_any_authenticated_packet_received(struct wg_peer *peer)
|
||||
{
|
||||
del_timer(&peer->timer_new_handshake);
|
||||
}
|
||||
|
||||
/* Should be called after a handshake initiation message is sent. */
|
||||
void wg_timers_handshake_initiated(struct wg_peer *peer)
|
||||
{
|
||||
mod_peer_timer(peer, &peer->timer_retransmit_handshake,
|
||||
jiffies + REKEY_TIMEOUT * HZ +
|
||||
prandom_u32_max(REKEY_TIMEOUT_JITTER_MAX_JIFFIES));
|
||||
}
|
||||
|
||||
/* Should be called after a handshake response message is received and processed
|
||||
* or when getting key confirmation via the first data message.
|
||||
*/
|
||||
void wg_timers_handshake_complete(struct wg_peer *peer)
|
||||
{
|
||||
del_timer(&peer->timer_retransmit_handshake);
|
||||
peer->timer_handshake_attempts = 0;
|
||||
peer->sent_lastminute_handshake = false;
|
||||
ktime_get_real_ts64(&peer->walltime_last_handshake);
|
||||
}
|
||||
|
||||
/* Should be called after an ephemeral key is created, which is before sending a
|
||||
* handshake response or after receiving a handshake response.
|
||||
*/
|
||||
void wg_timers_session_derived(struct wg_peer *peer)
|
||||
{
|
||||
mod_peer_timer(peer, &peer->timer_zero_key_material,
|
||||
jiffies + REJECT_AFTER_TIME * 3 * HZ);
|
||||
}
|
||||
|
||||
/* Should be called before a packet with authentication, whether
|
||||
* keepalive, data, or handshakem is sent, or after one is received.
|
||||
*/
|
||||
void wg_timers_any_authenticated_packet_traversal(struct wg_peer *peer)
|
||||
{
|
||||
if (peer->persistent_keepalive_interval)
|
||||
mod_peer_timer(peer, &peer->timer_persistent_keepalive,
|
||||
jiffies + peer->persistent_keepalive_interval * HZ);
|
||||
}
|
||||
|
||||
void wg_timers_init(struct wg_peer *peer)
|
||||
{
|
||||
timer_setup(&peer->timer_retransmit_handshake,
|
||||
wg_expired_retransmit_handshake, 0);
|
||||
timer_setup(&peer->timer_send_keepalive, wg_expired_send_keepalive, 0);
|
||||
timer_setup(&peer->timer_new_handshake, wg_expired_new_handshake, 0);
|
||||
timer_setup(&peer->timer_zero_key_material,
|
||||
wg_expired_zero_key_material, 0);
|
||||
timer_setup(&peer->timer_persistent_keepalive,
|
||||
wg_expired_send_persistent_keepalive, 0);
|
||||
INIT_WORK(&peer->clear_peer_work, wg_queued_expired_zero_key_material);
|
||||
peer->timer_handshake_attempts = 0;
|
||||
peer->sent_lastminute_handshake = false;
|
||||
peer->timer_need_another_keepalive = false;
|
||||
}
|
||||
|
||||
void wg_timers_stop(struct wg_peer *peer)
|
||||
{
|
||||
del_timer_sync(&peer->timer_retransmit_handshake);
|
||||
del_timer_sync(&peer->timer_send_keepalive);
|
||||
del_timer_sync(&peer->timer_new_handshake);
|
||||
del_timer_sync(&peer->timer_zero_key_material);
|
||||
del_timer_sync(&peer->timer_persistent_keepalive);
|
||||
flush_work(&peer->clear_peer_work);
|
||||
}
|
||||
31
drivers/net/wireguard/timers.h
Normal file
31
drivers/net/wireguard/timers.h
Normal file
|
|
@ -0,0 +1,31 @@
|
|||
/* SPDX-License-Identifier: GPL-2.0 */
|
||||
/*
|
||||
* Copyright (C) 2015-2019 Jason A. Donenfeld <Jason@zx2c4.com>. All Rights Reserved.
|
||||
*/
|
||||
|
||||
#ifndef _WG_TIMERS_H
|
||||
#define _WG_TIMERS_H
|
||||
|
||||
#include <linux/ktime.h>
|
||||
|
||||
struct wg_peer;
|
||||
|
||||
void wg_timers_init(struct wg_peer *peer);
|
||||
void wg_timers_stop(struct wg_peer *peer);
|
||||
void wg_timers_data_sent(struct wg_peer *peer);
|
||||
void wg_timers_data_received(struct wg_peer *peer);
|
||||
void wg_timers_any_authenticated_packet_sent(struct wg_peer *peer);
|
||||
void wg_timers_any_authenticated_packet_received(struct wg_peer *peer);
|
||||
void wg_timers_handshake_initiated(struct wg_peer *peer);
|
||||
void wg_timers_handshake_complete(struct wg_peer *peer);
|
||||
void wg_timers_session_derived(struct wg_peer *peer);
|
||||
void wg_timers_any_authenticated_packet_traversal(struct wg_peer *peer);
|
||||
|
||||
static inline bool wg_birthdate_has_expired(u64 birthday_nanoseconds,
|
||||
u64 expiration_seconds)
|
||||
{
|
||||
return (s64)(birthday_nanoseconds + expiration_seconds * NSEC_PER_SEC)
|
||||
<= (s64)ktime_get_coarse_boottime_ns();
|
||||
}
|
||||
|
||||
#endif /* _WG_TIMERS_H */
|
||||
1
drivers/net/wireguard/version.h
Normal file
1
drivers/net/wireguard/version.h
Normal file
|
|
@ -0,0 +1 @@
|
|||
#define WIREGUARD_VERSION "1.0.0"
|
||||
|
|
@ -15,9 +15,8 @@
|
|||
#ifndef _CRYPTO_CHACHA_H
|
||||
#define _CRYPTO_CHACHA_H
|
||||
|
||||
#include <crypto/skcipher.h>
|
||||
#include <asm/unaligned.h>
|
||||
#include <linux/types.h>
|
||||
#include <linux/crypto.h>
|
||||
|
||||
/* 32-bit stream position, then 96-bit nonce (RFC7539 convention) */
|
||||
#define CHACHA_IV_SIZE 16
|
||||
|
|
@ -26,30 +25,76 @@
|
|||
#define CHACHA_BLOCK_SIZE 64
|
||||
#define CHACHAPOLY_IV_SIZE 12
|
||||
|
||||
#define CHACHA_STATE_WORDS (CHACHA_BLOCK_SIZE / sizeof(u32))
|
||||
|
||||
/* 192-bit nonce, then 64-bit stream position */
|
||||
#define XCHACHA_IV_SIZE 32
|
||||
|
||||
struct chacha_ctx {
|
||||
u32 key[8];
|
||||
int nrounds;
|
||||
};
|
||||
|
||||
void chacha_block(u32 *state, u8 *stream, int nrounds);
|
||||
void chacha_block_generic(u32 *state, u8 *stream, int nrounds);
|
||||
static inline void chacha20_block(u32 *state, u8 *stream)
|
||||
{
|
||||
chacha_block(state, stream, 20);
|
||||
chacha_block_generic(state, stream, 20);
|
||||
}
|
||||
void hchacha_block(const u32 *in, u32 *out, int nrounds);
|
||||
|
||||
void crypto_chacha_init(u32 *state, const struct chacha_ctx *ctx, const u8 *iv);
|
||||
void hchacha_block_arch(const u32 *state, u32 *out, int nrounds);
|
||||
void hchacha_block_generic(const u32 *state, u32 *out, int nrounds);
|
||||
|
||||
int crypto_chacha20_setkey(struct crypto_skcipher *tfm, const u8 *key,
|
||||
unsigned int keysize);
|
||||
int crypto_chacha12_setkey(struct crypto_skcipher *tfm, const u8 *key,
|
||||
unsigned int keysize);
|
||||
static inline void hchacha_block(const u32 *state, u32 *out, int nrounds)
|
||||
{
|
||||
if (IS_ENABLED(CONFIG_CRYPTO_ARCH_HAVE_LIB_CHACHA))
|
||||
hchacha_block_arch(state, out, nrounds);
|
||||
else
|
||||
hchacha_block_generic(state, out, nrounds);
|
||||
}
|
||||
|
||||
int crypto_chacha_crypt(struct skcipher_request *req);
|
||||
int crypto_xchacha_crypt(struct skcipher_request *req);
|
||||
void chacha_init_arch(u32 *state, const u32 *key, const u8 *iv);
|
||||
static inline void chacha_init_generic(u32 *state, const u32 *key, const u8 *iv)
|
||||
{
|
||||
state[0] = 0x61707865; /* "expa" */
|
||||
state[1] = 0x3320646e; /* "nd 3" */
|
||||
state[2] = 0x79622d32; /* "2-by" */
|
||||
state[3] = 0x6b206574; /* "te k" */
|
||||
state[4] = key[0];
|
||||
state[5] = key[1];
|
||||
state[6] = key[2];
|
||||
state[7] = key[3];
|
||||
state[8] = key[4];
|
||||
state[9] = key[5];
|
||||
state[10] = key[6];
|
||||
state[11] = key[7];
|
||||
state[12] = get_unaligned_le32(iv + 0);
|
||||
state[13] = get_unaligned_le32(iv + 4);
|
||||
state[14] = get_unaligned_le32(iv + 8);
|
||||
state[15] = get_unaligned_le32(iv + 12);
|
||||
}
|
||||
|
||||
static inline void chacha_init(u32 *state, const u32 *key, const u8 *iv)
|
||||
{
|
||||
if (IS_ENABLED(CONFIG_CRYPTO_ARCH_HAVE_LIB_CHACHA))
|
||||
chacha_init_arch(state, key, iv);
|
||||
else
|
||||
chacha_init_generic(state, key, iv);
|
||||
}
|
||||
|
||||
void chacha_crypt_arch(u32 *state, u8 *dst, const u8 *src,
|
||||
unsigned int bytes, int nrounds);
|
||||
void chacha_crypt_generic(u32 *state, u8 *dst, const u8 *src,
|
||||
unsigned int bytes, int nrounds);
|
||||
|
||||
static inline void chacha_crypt(u32 *state, u8 *dst, const u8 *src,
|
||||
unsigned int bytes, int nrounds)
|
||||
{
|
||||
if (IS_ENABLED(CONFIG_CRYPTO_ARCH_HAVE_LIB_CHACHA))
|
||||
chacha_crypt_arch(state, dst, src, bytes, nrounds);
|
||||
else
|
||||
chacha_crypt_generic(state, dst, src, bytes, nrounds);
|
||||
}
|
||||
|
||||
static inline void chacha20_crypt(u32 *state, u8 *dst, const u8 *src,
|
||||
unsigned int bytes)
|
||||
{
|
||||
chacha_crypt(state, dst, src, bytes, 20);
|
||||
}
|
||||
|
||||
enum chacha_constants { /* expand 32-byte k */
|
||||
CHACHA_CONSTANT_EXPA = 0x61707865U,
|
||||
|
|
|
|||
50
include/crypto/chacha20poly1305.h
Normal file
50
include/crypto/chacha20poly1305.h
Normal file
|
|
@ -0,0 +1,50 @@
|
|||
/* SPDX-License-Identifier: GPL-2.0 OR MIT */
|
||||
/*
|
||||
* Copyright (C) 2015-2019 Jason A. Donenfeld <Jason@zx2c4.com>. All Rights Reserved.
|
||||
*/
|
||||
|
||||
#ifndef __CHACHA20POLY1305_H
|
||||
#define __CHACHA20POLY1305_H
|
||||
|
||||
#include <linux/types.h>
|
||||
#include <linux/scatterlist.h>
|
||||
|
||||
enum chacha20poly1305_lengths {
|
||||
XCHACHA20POLY1305_NONCE_SIZE = 24,
|
||||
CHACHA20POLY1305_KEY_SIZE = 32,
|
||||
CHACHA20POLY1305_AUTHTAG_SIZE = 16
|
||||
};
|
||||
|
||||
void chacha20poly1305_encrypt(u8 *dst, const u8 *src, const size_t src_len,
|
||||
const u8 *ad, const size_t ad_len,
|
||||
const u64 nonce,
|
||||
const u8 key[CHACHA20POLY1305_KEY_SIZE]);
|
||||
|
||||
bool __must_check
|
||||
chacha20poly1305_decrypt(u8 *dst, const u8 *src, const size_t src_len,
|
||||
const u8 *ad, const size_t ad_len, const u64 nonce,
|
||||
const u8 key[CHACHA20POLY1305_KEY_SIZE]);
|
||||
|
||||
void xchacha20poly1305_encrypt(u8 *dst, const u8 *src, const size_t src_len,
|
||||
const u8 *ad, const size_t ad_len,
|
||||
const u8 nonce[XCHACHA20POLY1305_NONCE_SIZE],
|
||||
const u8 key[CHACHA20POLY1305_KEY_SIZE]);
|
||||
|
||||
bool __must_check xchacha20poly1305_decrypt(
|
||||
u8 *dst, const u8 *src, const size_t src_len, const u8 *ad,
|
||||
const size_t ad_len, const u8 nonce[XCHACHA20POLY1305_NONCE_SIZE],
|
||||
const u8 key[CHACHA20POLY1305_KEY_SIZE]);
|
||||
|
||||
bool chacha20poly1305_encrypt_sg_inplace(struct scatterlist *src, size_t src_len,
|
||||
const u8 *ad, const size_t ad_len,
|
||||
const u64 nonce,
|
||||
const u8 key[CHACHA20POLY1305_KEY_SIZE]);
|
||||
|
||||
bool chacha20poly1305_decrypt_sg_inplace(struct scatterlist *src, size_t src_len,
|
||||
const u8 *ad, const size_t ad_len,
|
||||
const u64 nonce,
|
||||
const u8 key[CHACHA20POLY1305_KEY_SIZE]);
|
||||
|
||||
bool chacha20poly1305_selftest(void);
|
||||
|
||||
#endif /* __CHACHA20POLY1305_H */
|
||||
73
include/crypto/curve25519.h
Normal file
73
include/crypto/curve25519.h
Normal file
|
|
@ -0,0 +1,73 @@
|
|||
/* SPDX-License-Identifier: GPL-2.0 OR MIT */
|
||||
/*
|
||||
* Copyright (C) 2015-2019 Jason A. Donenfeld <Jason@zx2c4.com>. All Rights Reserved.
|
||||
*/
|
||||
|
||||
#ifndef CURVE25519_H
|
||||
#define CURVE25519_H
|
||||
|
||||
#include <crypto/algapi.h> // For crypto_memneq.
|
||||
#include <linux/types.h>
|
||||
#include <linux/random.h>
|
||||
|
||||
enum curve25519_lengths {
|
||||
CURVE25519_KEY_SIZE = 32
|
||||
};
|
||||
|
||||
extern const u8 curve25519_null_point[];
|
||||
extern const u8 curve25519_base_point[];
|
||||
|
||||
void curve25519_generic(u8 out[CURVE25519_KEY_SIZE],
|
||||
const u8 scalar[CURVE25519_KEY_SIZE],
|
||||
const u8 point[CURVE25519_KEY_SIZE]);
|
||||
|
||||
void curve25519_arch(u8 out[CURVE25519_KEY_SIZE],
|
||||
const u8 scalar[CURVE25519_KEY_SIZE],
|
||||
const u8 point[CURVE25519_KEY_SIZE]);
|
||||
|
||||
void curve25519_base_arch(u8 pub[CURVE25519_KEY_SIZE],
|
||||
const u8 secret[CURVE25519_KEY_SIZE]);
|
||||
|
||||
static inline
|
||||
bool __must_check curve25519(u8 mypublic[CURVE25519_KEY_SIZE],
|
||||
const u8 secret[CURVE25519_KEY_SIZE],
|
||||
const u8 basepoint[CURVE25519_KEY_SIZE])
|
||||
{
|
||||
if (IS_ENABLED(CONFIG_CRYPTO_ARCH_HAVE_LIB_CURVE25519) &&
|
||||
(!IS_ENABLED(CONFIG_CRYPTO_CURVE25519_X86) || IS_ENABLED(CONFIG_AS_ADX)))
|
||||
curve25519_arch(mypublic, secret, basepoint);
|
||||
else
|
||||
curve25519_generic(mypublic, secret, basepoint);
|
||||
return crypto_memneq(mypublic, curve25519_null_point,
|
||||
CURVE25519_KEY_SIZE);
|
||||
}
|
||||
|
||||
static inline bool
|
||||
__must_check curve25519_generate_public(u8 pub[CURVE25519_KEY_SIZE],
|
||||
const u8 secret[CURVE25519_KEY_SIZE])
|
||||
{
|
||||
if (unlikely(!crypto_memneq(secret, curve25519_null_point,
|
||||
CURVE25519_KEY_SIZE)))
|
||||
return false;
|
||||
|
||||
if (IS_ENABLED(CONFIG_CRYPTO_ARCH_HAVE_LIB_CURVE25519) &&
|
||||
(!IS_ENABLED(CONFIG_CRYPTO_CURVE25519_X86) || IS_ENABLED(CONFIG_AS_ADX)))
|
||||
curve25519_base_arch(pub, secret);
|
||||
else
|
||||
curve25519_generic(pub, secret, curve25519_base_point);
|
||||
return crypto_memneq(pub, curve25519_null_point, CURVE25519_KEY_SIZE);
|
||||
}
|
||||
|
||||
static inline void curve25519_clamp_secret(u8 secret[CURVE25519_KEY_SIZE])
|
||||
{
|
||||
secret[0] &= 248;
|
||||
secret[31] = (secret[31] & 127) | 64;
|
||||
}
|
||||
|
||||
static inline void curve25519_generate_secret(u8 secret[CURVE25519_KEY_SIZE])
|
||||
{
|
||||
get_random_bytes_wait(secret, CURVE25519_KEY_SIZE);
|
||||
curve25519_clamp_secret(secret);
|
||||
}
|
||||
|
||||
#endif /* CURVE25519_H */
|
||||
43
include/crypto/internal/chacha.h
Normal file
43
include/crypto/internal/chacha.h
Normal file
|
|
@ -0,0 +1,43 @@
|
|||
/* SPDX-License-Identifier: GPL-2.0 */
|
||||
|
||||
#ifndef _CRYPTO_INTERNAL_CHACHA_H
|
||||
#define _CRYPTO_INTERNAL_CHACHA_H
|
||||
|
||||
#include <crypto/chacha.h>
|
||||
#include <crypto/internal/skcipher.h>
|
||||
#include <linux/crypto.h>
|
||||
|
||||
struct chacha_ctx {
|
||||
u32 key[8];
|
||||
int nrounds;
|
||||
};
|
||||
|
||||
static inline int chacha_setkey(struct crypto_skcipher *tfm, const u8 *key,
|
||||
unsigned int keysize, int nrounds)
|
||||
{
|
||||
struct chacha_ctx *ctx = crypto_skcipher_ctx(tfm);
|
||||
int i;
|
||||
|
||||
if (keysize != CHACHA_KEY_SIZE)
|
||||
return -EINVAL;
|
||||
|
||||
for (i = 0; i < ARRAY_SIZE(ctx->key); i++)
|
||||
ctx->key[i] = get_unaligned_le32(key + i * sizeof(u32));
|
||||
|
||||
ctx->nrounds = nrounds;
|
||||
return 0;
|
||||
}
|
||||
|
||||
static inline int chacha20_setkey(struct crypto_skcipher *tfm, const u8 *key,
|
||||
unsigned int keysize)
|
||||
{
|
||||
return chacha_setkey(tfm, key, keysize, 20);
|
||||
}
|
||||
|
||||
static inline int chacha12_setkey(struct crypto_skcipher *tfm, const u8 *key,
|
||||
unsigned int keysize)
|
||||
{
|
||||
return chacha_setkey(tfm, key, keysize, 12);
|
||||
}
|
||||
|
||||
#endif /* _CRYPTO_CHACHA_H */
|
||||
34
include/crypto/internal/poly1305.h
Normal file
34
include/crypto/internal/poly1305.h
Normal file
|
|
@ -0,0 +1,34 @@
|
|||
/* SPDX-License-Identifier: GPL-2.0 */
|
||||
/*
|
||||
* Common values for the Poly1305 algorithm
|
||||
*/
|
||||
|
||||
#ifndef _CRYPTO_INTERNAL_POLY1305_H
|
||||
#define _CRYPTO_INTERNAL_POLY1305_H
|
||||
|
||||
#include <asm/unaligned.h>
|
||||
#include <linux/types.h>
|
||||
#include <crypto/poly1305.h>
|
||||
|
||||
/*
|
||||
* Poly1305 core functions. These only accept whole blocks; the caller must
|
||||
* handle any needed block buffering and padding. 'hibit' must be 1 for any
|
||||
* full blocks, or 0 for the final block if it had to be padded. If 'nonce' is
|
||||
* non-NULL, then it's added at the end to compute the Poly1305 MAC. Otherwise,
|
||||
* only the ε-almost-∆-universal hash function (not the full MAC) is computed.
|
||||
*/
|
||||
|
||||
void poly1305_core_setkey(struct poly1305_core_key *key,
|
||||
const u8 raw_key[POLY1305_BLOCK_SIZE]);
|
||||
static inline void poly1305_core_init(struct poly1305_state *state)
|
||||
{
|
||||
*state = (struct poly1305_state){};
|
||||
}
|
||||
|
||||
void poly1305_core_blocks(struct poly1305_state *state,
|
||||
const struct poly1305_core_key *key, const void *src,
|
||||
unsigned int nblocks, u32 hibit);
|
||||
void poly1305_core_emit(const struct poly1305_state *state, const u32 nonce[4],
|
||||
void *dst);
|
||||
|
||||
#endif
|
||||
|
|
@ -7,7 +7,7 @@
|
|||
#define _NHPOLY1305_H
|
||||
|
||||
#include <crypto/hash.h>
|
||||
#include <crypto/poly1305.h>
|
||||
#include <crypto/internal/poly1305.h>
|
||||
|
||||
/* NH parameterization: */
|
||||
|
||||
|
|
@ -33,7 +33,7 @@
|
|||
#define NHPOLY1305_KEY_SIZE (POLY1305_BLOCK_SIZE + NH_KEY_BYTES)
|
||||
|
||||
struct nhpoly1305_key {
|
||||
struct poly1305_key poly_key;
|
||||
struct poly1305_core_key poly_key;
|
||||
u32 nh_key[NH_KEY_WORDS];
|
||||
};
|
||||
|
||||
|
|
|
|||
|
|
@ -13,52 +13,87 @@
|
|||
#define POLY1305_KEY_SIZE 32
|
||||
#define POLY1305_DIGEST_SIZE 16
|
||||
|
||||
/* The poly1305_key and poly1305_state types are mostly opaque and
|
||||
* implementation-defined. Limbs might be in base 2^64 or base 2^26, or
|
||||
* different yet. The union type provided keeps these 64-bit aligned for the
|
||||
* case in which this is implemented using 64x64 multiplies.
|
||||
*/
|
||||
|
||||
struct poly1305_key {
|
||||
u32 r[5]; /* key, base 2^26 */
|
||||
union {
|
||||
u32 r[5];
|
||||
u64 r64[3];
|
||||
};
|
||||
};
|
||||
|
||||
struct poly1305_core_key {
|
||||
struct poly1305_key key;
|
||||
struct poly1305_key precomputed_s;
|
||||
};
|
||||
|
||||
struct poly1305_state {
|
||||
u32 h[5]; /* accumulator, base 2^26 */
|
||||
union {
|
||||
u32 h[5];
|
||||
u64 h64[3];
|
||||
};
|
||||
};
|
||||
|
||||
struct poly1305_desc_ctx {
|
||||
/* key */
|
||||
struct poly1305_key r;
|
||||
/* finalize key */
|
||||
u32 s[4];
|
||||
/* accumulator */
|
||||
struct poly1305_state h;
|
||||
/* partial buffer */
|
||||
u8 buf[POLY1305_BLOCK_SIZE];
|
||||
/* bytes used in partial buffer */
|
||||
unsigned int buflen;
|
||||
/* r key has been set */
|
||||
bool rset;
|
||||
/* s key has been set */
|
||||
/* how many keys have been set in r[] */
|
||||
unsigned short rset;
|
||||
/* whether s[] has been set */
|
||||
bool sset;
|
||||
/* finalize key */
|
||||
u32 s[4];
|
||||
/* accumulator */
|
||||
struct poly1305_state h;
|
||||
/* key */
|
||||
union {
|
||||
struct poly1305_key opaque_r[CONFIG_CRYPTO_LIB_POLY1305_RSIZE];
|
||||
struct poly1305_core_key core_r;
|
||||
};
|
||||
};
|
||||
|
||||
/*
|
||||
* Poly1305 core functions. These implement the ε-almost-∆-universal hash
|
||||
* function underlying the Poly1305 MAC, i.e. they don't add an encrypted nonce
|
||||
* ("s key") at the end. They also only support block-aligned inputs.
|
||||
*/
|
||||
void poly1305_core_setkey(struct poly1305_key *key, const u8 *raw_key);
|
||||
static inline void poly1305_core_init(struct poly1305_state *state)
|
||||
{
|
||||
memset(state->h, 0, sizeof(state->h));
|
||||
}
|
||||
void poly1305_core_blocks(struct poly1305_state *state,
|
||||
const struct poly1305_key *key,
|
||||
const void *src, unsigned int nblocks);
|
||||
void poly1305_core_emit(const struct poly1305_state *state, void *dst);
|
||||
void poly1305_init_arch(struct poly1305_desc_ctx *desc,
|
||||
const u8 key[POLY1305_KEY_SIZE]);
|
||||
void poly1305_init_generic(struct poly1305_desc_ctx *desc,
|
||||
const u8 key[POLY1305_KEY_SIZE]);
|
||||
|
||||
/* Crypto API helper functions for the Poly1305 MAC */
|
||||
int crypto_poly1305_init(struct shash_desc *desc);
|
||||
unsigned int crypto_poly1305_setdesckey(struct poly1305_desc_ctx *dctx,
|
||||
const u8 *src, unsigned int srclen);
|
||||
int crypto_poly1305_update(struct shash_desc *desc,
|
||||
const u8 *src, unsigned int srclen);
|
||||
int crypto_poly1305_final(struct shash_desc *desc, u8 *dst);
|
||||
static inline void poly1305_init(struct poly1305_desc_ctx *desc, const u8 *key)
|
||||
{
|
||||
if (IS_ENABLED(CONFIG_CRYPTO_ARCH_HAVE_LIB_POLY1305))
|
||||
poly1305_init_arch(desc, key);
|
||||
else
|
||||
poly1305_init_generic(desc, key);
|
||||
}
|
||||
|
||||
void poly1305_update_arch(struct poly1305_desc_ctx *desc, const u8 *src,
|
||||
unsigned int nbytes);
|
||||
void poly1305_update_generic(struct poly1305_desc_ctx *desc, const u8 *src,
|
||||
unsigned int nbytes);
|
||||
|
||||
static inline void poly1305_update(struct poly1305_desc_ctx *desc,
|
||||
const u8 *src, unsigned int nbytes)
|
||||
{
|
||||
if (IS_ENABLED(CONFIG_CRYPTO_ARCH_HAVE_LIB_POLY1305))
|
||||
poly1305_update_arch(desc, src, nbytes);
|
||||
else
|
||||
poly1305_update_generic(desc, src, nbytes);
|
||||
}
|
||||
|
||||
void poly1305_final_arch(struct poly1305_desc_ctx *desc, u8 *digest);
|
||||
void poly1305_final_generic(struct poly1305_desc_ctx *desc, u8 *digest);
|
||||
|
||||
static inline void poly1305_final(struct poly1305_desc_ctx *desc, u8 *digest)
|
||||
{
|
||||
if (IS_ENABLED(CONFIG_CRYPTO_ARCH_HAVE_LIB_POLY1305))
|
||||
poly1305_final_arch(desc, digest);
|
||||
else
|
||||
poly1305_final_generic(desc, digest);
|
||||
}
|
||||
|
||||
#endif
|
||||
|
|
|
|||
|
|
@ -123,7 +123,7 @@ extern int crypto_sha512_finup(struct shash_desc *desc, const u8 *data,
|
|||
* For details see lib/crypto/sha256.c
|
||||
*/
|
||||
|
||||
static inline int sha256_init(struct sha256_state *sctx)
|
||||
static inline void sha256_init(struct sha256_state *sctx)
|
||||
{
|
||||
sctx->state[0] = SHA256_H0;
|
||||
sctx->state[1] = SHA256_H1;
|
||||
|
|
@ -134,14 +134,12 @@ static inline int sha256_init(struct sha256_state *sctx)
|
|||
sctx->state[6] = SHA256_H6;
|
||||
sctx->state[7] = SHA256_H7;
|
||||
sctx->count = 0;
|
||||
|
||||
return 0;
|
||||
}
|
||||
extern int sha256_update(struct sha256_state *sctx, const u8 *input,
|
||||
unsigned int length);
|
||||
extern int sha256_final(struct sha256_state *sctx, u8 *hash);
|
||||
void sha256_update(struct sha256_state *sctx, const u8 *data, unsigned int len);
|
||||
void sha256_final(struct sha256_state *sctx, u8 *out);
|
||||
void sha256(const u8 *data, unsigned int len, u8 *out);
|
||||
|
||||
static inline int sha224_init(struct sha256_state *sctx)
|
||||
static inline void sha224_init(struct sha256_state *sctx)
|
||||
{
|
||||
sctx->state[0] = SHA224_H0;
|
||||
sctx->state[1] = SHA224_H1;
|
||||
|
|
@ -152,11 +150,8 @@ static inline int sha224_init(struct sha256_state *sctx)
|
|||
sctx->state[6] = SHA224_H6;
|
||||
sctx->state[7] = SHA224_H7;
|
||||
sctx->count = 0;
|
||||
|
||||
return 0;
|
||||
}
|
||||
extern int sha224_update(struct sha256_state *sctx, const u8 *input,
|
||||
unsigned int length);
|
||||
extern int sha224_final(struct sha256_state *sctx, u8 *hash);
|
||||
void sha224_update(struct sha256_state *sctx, const u8 *data, unsigned int len);
|
||||
void sha224_final(struct sha256_state *sctx, u8 *out);
|
||||
|
||||
#endif
|
||||
|
|
|
|||
|
|
@ -22,14 +22,16 @@ static inline int sha224_base_init(struct shash_desc *desc)
|
|||
{
|
||||
struct sha256_state *sctx = shash_desc_ctx(desc);
|
||||
|
||||
return sha224_init(sctx);
|
||||
sha224_init(sctx);
|
||||
return 0;
|
||||
}
|
||||
|
||||
static inline int sha256_base_init(struct shash_desc *desc)
|
||||
{
|
||||
struct sha256_state *sctx = shash_desc_ctx(desc);
|
||||
|
||||
return sha256_init(sctx);
|
||||
sha256_init(sctx);
|
||||
return 0;
|
||||
}
|
||||
|
||||
static inline int sha256_base_do_update(struct shash_desc *desc,
|
||||
|
|
|
|||
|
|
@ -79,6 +79,17 @@ static inline void dst_cache_reset(struct dst_cache *dst_cache)
|
|||
dst_cache->reset_ts = jiffies;
|
||||
}
|
||||
|
||||
/**
|
||||
* dst_cache_reset_now - invalidate the cache contents immediately
|
||||
* @dst_cache: the cache
|
||||
*
|
||||
* The caller must be sure there are no concurrent users, as this frees
|
||||
* all dst_cache users immediately, rather than waiting for the next
|
||||
* per-cpu usage like dst_cache_reset does. Most callers should use the
|
||||
* higher speed lazily-freed dst_cache_reset function instead.
|
||||
*/
|
||||
void dst_cache_reset_now(struct dst_cache *dst_cache);
|
||||
|
||||
/**
|
||||
* dst_cache_init - initialize the cache, allocating the required storage
|
||||
* @dst_cache: the cache
|
||||
|
|
|
|||
|
|
@ -289,6 +289,9 @@ int ip_tunnel_newlink(struct net_device *dev, struct nlattr *tb[],
|
|||
struct ip_tunnel_parm *p, __u32 fwmark);
|
||||
void ip_tunnel_setup(struct net_device *dev, unsigned int net_id);
|
||||
|
||||
extern const struct header_ops ip_tunnel_header_ops;
|
||||
__be16 ip_tunnel_parse_protocol(const struct sk_buff *skb);
|
||||
|
||||
struct ip_tunnel_encap_ops {
|
||||
size_t (*encap_hlen)(struct ip_tunnel_encap *e);
|
||||
int (*build_header)(struct sk_buff *skb, struct ip_tunnel_encap *e,
|
||||
|
|
|
|||
196
include/uapi/linux/wireguard.h
Normal file
196
include/uapi/linux/wireguard.h
Normal file
|
|
@ -0,0 +1,196 @@
|
|||
/* SPDX-License-Identifier: (GPL-2.0 WITH Linux-syscall-note) OR MIT */
|
||||
/*
|
||||
* Copyright (C) 2015-2019 Jason A. Donenfeld <Jason@zx2c4.com>. All Rights Reserved.
|
||||
*
|
||||
* Documentation
|
||||
* =============
|
||||
*
|
||||
* The below enums and macros are for interfacing with WireGuard, using generic
|
||||
* netlink, with family WG_GENL_NAME and version WG_GENL_VERSION. It defines two
|
||||
* methods: get and set. Note that while they share many common attributes,
|
||||
* these two functions actually accept a slightly different set of inputs and
|
||||
* outputs.
|
||||
*
|
||||
* WG_CMD_GET_DEVICE
|
||||
* -----------------
|
||||
*
|
||||
* May only be called via NLM_F_REQUEST | NLM_F_DUMP. The command should contain
|
||||
* one but not both of:
|
||||
*
|
||||
* WGDEVICE_A_IFINDEX: NLA_U32
|
||||
* WGDEVICE_A_IFNAME: NLA_NUL_STRING, maxlen IFNAMSIZ - 1
|
||||
*
|
||||
* The kernel will then return several messages (NLM_F_MULTI) containing the
|
||||
* following tree of nested items:
|
||||
*
|
||||
* WGDEVICE_A_IFINDEX: NLA_U32
|
||||
* WGDEVICE_A_IFNAME: NLA_NUL_STRING, maxlen IFNAMSIZ - 1
|
||||
* WGDEVICE_A_PRIVATE_KEY: NLA_EXACT_LEN, len WG_KEY_LEN
|
||||
* WGDEVICE_A_PUBLIC_KEY: NLA_EXACT_LEN, len WG_KEY_LEN
|
||||
* WGDEVICE_A_LISTEN_PORT: NLA_U16
|
||||
* WGDEVICE_A_FWMARK: NLA_U32
|
||||
* WGDEVICE_A_PEERS: NLA_NESTED
|
||||
* 0: NLA_NESTED
|
||||
* WGPEER_A_PUBLIC_KEY: NLA_EXACT_LEN, len WG_KEY_LEN
|
||||
* WGPEER_A_PRESHARED_KEY: NLA_EXACT_LEN, len WG_KEY_LEN
|
||||
* WGPEER_A_ENDPOINT: NLA_MIN_LEN(struct sockaddr), struct sockaddr_in or struct sockaddr_in6
|
||||
* WGPEER_A_PERSISTENT_KEEPALIVE_INTERVAL: NLA_U16
|
||||
* WGPEER_A_LAST_HANDSHAKE_TIME: NLA_EXACT_LEN, struct __kernel_timespec
|
||||
* WGPEER_A_RX_BYTES: NLA_U64
|
||||
* WGPEER_A_TX_BYTES: NLA_U64
|
||||
* WGPEER_A_ALLOWEDIPS: NLA_NESTED
|
||||
* 0: NLA_NESTED
|
||||
* WGALLOWEDIP_A_FAMILY: NLA_U16
|
||||
* WGALLOWEDIP_A_IPADDR: NLA_MIN_LEN(struct in_addr), struct in_addr or struct in6_addr
|
||||
* WGALLOWEDIP_A_CIDR_MASK: NLA_U8
|
||||
* 0: NLA_NESTED
|
||||
* ...
|
||||
* 0: NLA_NESTED
|
||||
* ...
|
||||
* ...
|
||||
* WGPEER_A_PROTOCOL_VERSION: NLA_U32
|
||||
* 0: NLA_NESTED
|
||||
* ...
|
||||
* ...
|
||||
*
|
||||
* It is possible that all of the allowed IPs of a single peer will not
|
||||
* fit within a single netlink message. In that case, the same peer will
|
||||
* be written in the following message, except it will only contain
|
||||
* WGPEER_A_PUBLIC_KEY and WGPEER_A_ALLOWEDIPS. This may occur several
|
||||
* times in a row for the same peer. It is then up to the receiver to
|
||||
* coalesce adjacent peers. Likewise, it is possible that all peers will
|
||||
* not fit within a single message. So, subsequent peers will be sent
|
||||
* in following messages, except those will only contain WGDEVICE_A_IFNAME
|
||||
* and WGDEVICE_A_PEERS. It is then up to the receiver to coalesce these
|
||||
* messages to form the complete list of peers.
|
||||
*
|
||||
* Since this is an NLA_F_DUMP command, the final message will always be
|
||||
* NLMSG_DONE, even if an error occurs. However, this NLMSG_DONE message
|
||||
* contains an integer error code. It is either zero or a negative error
|
||||
* code corresponding to the errno.
|
||||
*
|
||||
* WG_CMD_SET_DEVICE
|
||||
* -----------------
|
||||
*
|
||||
* May only be called via NLM_F_REQUEST. The command should contain the
|
||||
* following tree of nested items, containing one but not both of
|
||||
* WGDEVICE_A_IFINDEX and WGDEVICE_A_IFNAME:
|
||||
*
|
||||
* WGDEVICE_A_IFINDEX: NLA_U32
|
||||
* WGDEVICE_A_IFNAME: NLA_NUL_STRING, maxlen IFNAMSIZ - 1
|
||||
* WGDEVICE_A_FLAGS: NLA_U32, 0 or WGDEVICE_F_REPLACE_PEERS if all current
|
||||
* peers should be removed prior to adding the list below.
|
||||
* WGDEVICE_A_PRIVATE_KEY: len WG_KEY_LEN, all zeros to remove
|
||||
* WGDEVICE_A_LISTEN_PORT: NLA_U16, 0 to choose randomly
|
||||
* WGDEVICE_A_FWMARK: NLA_U32, 0 to disable
|
||||
* WGDEVICE_A_PEERS: NLA_NESTED
|
||||
* 0: NLA_NESTED
|
||||
* WGPEER_A_PUBLIC_KEY: len WG_KEY_LEN
|
||||
* WGPEER_A_FLAGS: NLA_U32, 0 and/or WGPEER_F_REMOVE_ME if the
|
||||
* specified peer should not exist at the end of the
|
||||
* operation, rather than added/updated and/or
|
||||
* WGPEER_F_REPLACE_ALLOWEDIPS if all current allowed
|
||||
* IPs of this peer should be removed prior to adding
|
||||
* the list below and/or WGPEER_F_UPDATE_ONLY if the
|
||||
* peer should only be set if it already exists.
|
||||
* WGPEER_A_PRESHARED_KEY: len WG_KEY_LEN, all zeros to remove
|
||||
* WGPEER_A_ENDPOINT: struct sockaddr_in or struct sockaddr_in6
|
||||
* WGPEER_A_PERSISTENT_KEEPALIVE_INTERVAL: NLA_U16, 0 to disable
|
||||
* WGPEER_A_ALLOWEDIPS: NLA_NESTED
|
||||
* 0: NLA_NESTED
|
||||
* WGALLOWEDIP_A_FAMILY: NLA_U16
|
||||
* WGALLOWEDIP_A_IPADDR: struct in_addr or struct in6_addr
|
||||
* WGALLOWEDIP_A_CIDR_MASK: NLA_U8
|
||||
* 0: NLA_NESTED
|
||||
* ...
|
||||
* 0: NLA_NESTED
|
||||
* ...
|
||||
* ...
|
||||
* WGPEER_A_PROTOCOL_VERSION: NLA_U32, should not be set or used at
|
||||
* all by most users of this API, as the
|
||||
* most recent protocol will be used when
|
||||
* this is unset. Otherwise, must be set
|
||||
* to 1.
|
||||
* 0: NLA_NESTED
|
||||
* ...
|
||||
* ...
|
||||
*
|
||||
* It is possible that the amount of configuration data exceeds that of
|
||||
* the maximum message length accepted by the kernel. In that case, several
|
||||
* messages should be sent one after another, with each successive one
|
||||
* filling in information not contained in the prior. Note that if
|
||||
* WGDEVICE_F_REPLACE_PEERS is specified in the first message, it probably
|
||||
* should not be specified in fragments that come after, so that the list
|
||||
* of peers is only cleared the first time but appended after. Likewise for
|
||||
* peers, if WGPEER_F_REPLACE_ALLOWEDIPS is specified in the first message
|
||||
* of a peer, it likely should not be specified in subsequent fragments.
|
||||
*
|
||||
* If an error occurs, NLMSG_ERROR will reply containing an errno.
|
||||
*/
|
||||
|
||||
#ifndef _WG_UAPI_WIREGUARD_H
|
||||
#define _WG_UAPI_WIREGUARD_H
|
||||
|
||||
#define WG_GENL_NAME "wireguard"
|
||||
#define WG_GENL_VERSION 1
|
||||
|
||||
#define WG_KEY_LEN 32
|
||||
|
||||
enum wg_cmd {
|
||||
WG_CMD_GET_DEVICE,
|
||||
WG_CMD_SET_DEVICE,
|
||||
__WG_CMD_MAX
|
||||
};
|
||||
#define WG_CMD_MAX (__WG_CMD_MAX - 1)
|
||||
|
||||
enum wgdevice_flag {
|
||||
WGDEVICE_F_REPLACE_PEERS = 1U << 0,
|
||||
__WGDEVICE_F_ALL = WGDEVICE_F_REPLACE_PEERS
|
||||
};
|
||||
enum wgdevice_attribute {
|
||||
WGDEVICE_A_UNSPEC,
|
||||
WGDEVICE_A_IFINDEX,
|
||||
WGDEVICE_A_IFNAME,
|
||||
WGDEVICE_A_PRIVATE_KEY,
|
||||
WGDEVICE_A_PUBLIC_KEY,
|
||||
WGDEVICE_A_FLAGS,
|
||||
WGDEVICE_A_LISTEN_PORT,
|
||||
WGDEVICE_A_FWMARK,
|
||||
WGDEVICE_A_PEERS,
|
||||
__WGDEVICE_A_LAST
|
||||
};
|
||||
#define WGDEVICE_A_MAX (__WGDEVICE_A_LAST - 1)
|
||||
|
||||
enum wgpeer_flag {
|
||||
WGPEER_F_REMOVE_ME = 1U << 0,
|
||||
WGPEER_F_REPLACE_ALLOWEDIPS = 1U << 1,
|
||||
WGPEER_F_UPDATE_ONLY = 1U << 2,
|
||||
__WGPEER_F_ALL = WGPEER_F_REMOVE_ME | WGPEER_F_REPLACE_ALLOWEDIPS |
|
||||
WGPEER_F_UPDATE_ONLY
|
||||
};
|
||||
enum wgpeer_attribute {
|
||||
WGPEER_A_UNSPEC,
|
||||
WGPEER_A_PUBLIC_KEY,
|
||||
WGPEER_A_PRESHARED_KEY,
|
||||
WGPEER_A_FLAGS,
|
||||
WGPEER_A_ENDPOINT,
|
||||
WGPEER_A_PERSISTENT_KEEPALIVE_INTERVAL,
|
||||
WGPEER_A_LAST_HANDSHAKE_TIME,
|
||||
WGPEER_A_RX_BYTES,
|
||||
WGPEER_A_TX_BYTES,
|
||||
WGPEER_A_ALLOWEDIPS,
|
||||
WGPEER_A_PROTOCOL_VERSION,
|
||||
__WGPEER_A_LAST
|
||||
};
|
||||
#define WGPEER_A_MAX (__WGPEER_A_LAST - 1)
|
||||
|
||||
enum wgallowedip_attribute {
|
||||
WGALLOWEDIP_A_UNSPEC,
|
||||
WGALLOWEDIP_A_FAMILY,
|
||||
WGALLOWEDIP_A_IPADDR,
|
||||
WGALLOWEDIP_A_CIDR_MASK,
|
||||
__WGALLOWEDIP_A_LAST
|
||||
};
|
||||
#define WGALLOWEDIP_A_MAX (__WGALLOWEDIP_A_LAST - 1)
|
||||
|
||||
#endif /* _WG_UAPI_WIREGUARD_H */
|
||||
|
|
@ -2247,6 +2247,7 @@ schedule_hrtimeout_range_clock(ktime_t *expires, u64 delta,
|
|||
|
||||
return !t.task ? 0 : -EINTR;
|
||||
}
|
||||
EXPORT_SYMBOL_GPL(schedule_hrtimeout_range_clock);
|
||||
|
||||
/**
|
||||
* schedule_hrtimeout_range - sleep until timeout
|
||||
|
|
|
|||
|
|
@ -26,8 +26,7 @@ endif
|
|||
|
||||
lib-y := ctype.o string.o vsprintf.o cmdline.o \
|
||||
rbtree.o radix-tree.o timerqueue.o xarray.o \
|
||||
idr.o extable.o \
|
||||
sha1.o chacha.o irq_regs.o argv_split.o \
|
||||
idr.o extable.o sha1.o irq_regs.o argv_split.o \
|
||||
flex_proportions.o ratelimit.o show_mem.o \
|
||||
is_single_threaded.o plist.o decompress.o kobject_uevent.o \
|
||||
earlycpio.o seq_buf.o siphash.o dec_and_lock.o \
|
||||
|
|
|
|||
|
|
@ -24,9 +24,99 @@ config CRYPTO_LIB_BLAKE2S_GENERIC
|
|||
implementation is enabled, this implementation serves the users
|
||||
of CRYPTO_LIB_BLAKE2S.
|
||||
|
||||
config CRYPTO_ARCH_HAVE_LIB_CHACHA
|
||||
tristate
|
||||
help
|
||||
Declares whether the architecture provides an arch-specific
|
||||
accelerated implementation of the ChaCha library interface,
|
||||
either builtin or as a module.
|
||||
|
||||
config CRYPTO_LIB_CHACHA_GENERIC
|
||||
tristate
|
||||
select CRYPTO_ALGAPI
|
||||
help
|
||||
This symbol can be depended upon by arch implementations of the
|
||||
ChaCha library interface that require the generic code as a
|
||||
fallback, e.g., for SIMD implementations. If no arch specific
|
||||
implementation is enabled, this implementation serves the users
|
||||
of CRYPTO_LIB_CHACHA.
|
||||
|
||||
config CRYPTO_LIB_CHACHA
|
||||
tristate "ChaCha library interface"
|
||||
depends on CRYPTO_ARCH_HAVE_LIB_CHACHA || !CRYPTO_ARCH_HAVE_LIB_CHACHA
|
||||
select CRYPTO_LIB_CHACHA_GENERIC if CRYPTO_ARCH_HAVE_LIB_CHACHA=n
|
||||
help
|
||||
Enable the ChaCha library interface. This interface may be fulfilled
|
||||
by either the generic implementation or an arch-specific one, if one
|
||||
is available and enabled.
|
||||
|
||||
config CRYPTO_ARCH_HAVE_LIB_CURVE25519
|
||||
tristate
|
||||
help
|
||||
Declares whether the architecture provides an arch-specific
|
||||
accelerated implementation of the Curve25519 library interface,
|
||||
either builtin or as a module.
|
||||
|
||||
config CRYPTO_LIB_CURVE25519_GENERIC
|
||||
tristate
|
||||
help
|
||||
This symbol can be depended upon by arch implementations of the
|
||||
Curve25519 library interface that require the generic code as a
|
||||
fallback, e.g., for SIMD implementations. If no arch specific
|
||||
implementation is enabled, this implementation serves the users
|
||||
of CRYPTO_LIB_CURVE25519.
|
||||
|
||||
config CRYPTO_LIB_CURVE25519
|
||||
tristate "Curve25519 scalar multiplication library"
|
||||
depends on CRYPTO_ARCH_HAVE_LIB_CURVE25519 || !CRYPTO_ARCH_HAVE_LIB_CURVE25519
|
||||
select CRYPTO_LIB_CURVE25519_GENERIC if CRYPTO_ARCH_HAVE_LIB_CURVE25519=n
|
||||
help
|
||||
Enable the Curve25519 library interface. This interface may be
|
||||
fulfilled by either the generic implementation or an arch-specific
|
||||
one, if one is available and enabled.
|
||||
|
||||
config CRYPTO_LIB_DES
|
||||
tristate
|
||||
|
||||
config CRYPTO_LIB_POLY1305_RSIZE
|
||||
int
|
||||
default 2 if MIPS
|
||||
default 11 if X86_64
|
||||
default 9 if ARM || ARM64
|
||||
default 1
|
||||
|
||||
config CRYPTO_ARCH_HAVE_LIB_POLY1305
|
||||
tristate
|
||||
help
|
||||
Declares whether the architecture provides an arch-specific
|
||||
accelerated implementation of the Poly1305 library interface,
|
||||
either builtin or as a module.
|
||||
|
||||
config CRYPTO_LIB_POLY1305_GENERIC
|
||||
tristate
|
||||
help
|
||||
This symbol can be depended upon by arch implementations of the
|
||||
Poly1305 library interface that require the generic code as a
|
||||
fallback, e.g., for SIMD implementations. If no arch specific
|
||||
implementation is enabled, this implementation serves the users
|
||||
of CRYPTO_LIB_POLY1305.
|
||||
|
||||
config CRYPTO_LIB_POLY1305
|
||||
tristate "Poly1305 library interface"
|
||||
depends on CRYPTO_ARCH_HAVE_LIB_POLY1305 || !CRYPTO_ARCH_HAVE_LIB_POLY1305
|
||||
select CRYPTO_LIB_POLY1305_GENERIC if CRYPTO_ARCH_HAVE_LIB_POLY1305=n
|
||||
help
|
||||
Enable the Poly1305 library interface. This interface may be fulfilled
|
||||
by either the generic implementation or an arch-specific one, if one
|
||||
is available and enabled.
|
||||
|
||||
config CRYPTO_LIB_CHACHA20POLY1305
|
||||
tristate "ChaCha20-Poly1305 AEAD support (8-byte nonce library version)"
|
||||
depends on CRYPTO_ARCH_HAVE_LIB_CHACHA || !CRYPTO_ARCH_HAVE_LIB_CHACHA
|
||||
depends on CRYPTO_ARCH_HAVE_LIB_POLY1305 || !CRYPTO_ARCH_HAVE_LIB_POLY1305
|
||||
select CRYPTO_LIB_CHACHA
|
||||
select CRYPTO_LIB_POLY1305
|
||||
|
||||
config CRYPTO_LIB_SHA256
|
||||
tristate
|
||||
|
||||
|
|
|
|||
Some files were not shown because too many files have changed in this diff Show more
Loading…
Reference in a new issue