Skip to content

silk: Add Arm NEON silk_VAD_GetSA_Q8 - #483

Open
czoli1976 wants to merge 2 commits into
xiph:mainfrom
czoli1976:silk-arm-vad
Open

czoli1976 wants to merge 2 commits into
xiph:mainfrom
czoli1976:silk-arm-vad

Conversation

@czoli1976

Copy link
Copy Markdown

Summary

silk_VAD_GetSA_Q8 had an x86 SSE4.1 implementation but no Arm one, even though it runs on every SILK/hybrid frame in the default (float) build. This adds a NEON version, mirroring the SSE4.1 one.

What's vectorised

The per-subframe energy sum-of-squares — (X[i] >> 3)^2 accumulated in int32 — 8 samples per iteration via vshrq_n_s16 + paired vmlal_s16 (low/high), with a scalar tail and a horizontal vaddvq_s32. Everything else (analysis filterbank, noise estimation, SNR/tilt) is identical to the C reference, exactly as the SSE4.1 version does.

  • Bit-exact with silk_VAD_GetSA_Q8_c (exact integer sum of squares, no overflow), validated by the existing OPUS_CHECK_ASM full-encoder-state memcmp.
  • As on x86, silk_VAD_GetNoiseLevels becomes exported (instead of static inline in VAD.c) when NEON is enabled, so the kernel can call it.

Dispatch / wiring

Uses the existing OVERRIDE_silk_VAD_GetSA_Q8 hook: a new silk/arm/VAD_arm.h provides the PRESUME (direct call) and RTCD (SILK_VAD_GETSA_Q8_IMPL table in arm_silk_map.c) dispatch, mirroring silk/x86/main_sse.h. The source is added to the common SILK_SOURCES_ARM_NEON_INTR group, which is already wired in autotools / CMake / Meson — so no build-system changes are needed.

Numbers (Apple M4)

  • Microbench of the vectorised loop at the real subband lengths (10–80 samples): ~1.1–1.7× over scalar.
  • End-to-end within run-to-run noise (VAD is a small per-frame cost), bitstream unchanged (bit-exact).
  • Full meson test suite passes.

This is the last of the silk-side x86-has-it/ARM-doesn't parity gaps that runs in the default build.

silk_VAD_GetSA_Q8 had an x86 SSE4.1 implementation but no Arm one, and it
runs on every SILK/hybrid frame in the default (float) build. Add a NEON
version mirroring the SSE4.1 one: it vectorises the per-subframe energy
sum-of-squares ((X[i] >> 3)^2 accumulated in int32), 8 samples per iteration
via vshrq_n_s16 + paired vmlal_s16, with a scalar tail. Bit-exact with the C
reference (exact integer sum, no overflow), validated by the existing
OPUS_CHECK_ASM full-state memcmp.

As on x86, silk_VAD_GetNoiseLevels is exported (rather than static inline in
VAD.c) when NEON is enabled so the kernel can call it. Dispatched via the
existing OVERRIDE_silk_VAD_GetSA_Q8 hook (PRESUME + an RTCD table in
arm_silk_map.c); the source goes in the common SILK_SOURCES_ARM_NEON_INTR
group, already wired in autotools/CMake/Meson.

Microbench on Apple M4 (the real subband lengths, 10-80): ~1.1-1.7x over
scalar; E2E within run-to-run noise (VAD is a small per-frame cost). Full
meson test suite passes.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@lpi

lpi commented Sep 16, 2026 •

Copy link
Copy Markdown

Tested head: 0f4f92500c7af1acc3caf9f68b54328dc281c70b.

With #483 applied on top of #503, the approximate combined production-c10 encoder improvement versus stock is +4.0159% thread CPU / +4.0195% wall. Positive means faster/lower per-call cost; negative means slower. These are compounded cross-run estimates from the direct #503-versus-stock result and the incremental #483-versus-#503 result, not a direct stock arm.

#483 itself measured -0.0578% CPU / -0.0658% wall versus #503, so it did not improve encoding over #503. Route instrumentation confirmed that the modified VAD path executed in encode/re-encode, and the output/state checks were exact.

Separately, the unguarded AArch64-only vaddvq_s32 breaks ARMv7 compilation and should be guarded if ARMv7 remains supported.

vaddvq_s32 is AArch64-only and breaks ARMv7 compilation of the
silk_VAD_GetSA_Q8_neon kernel (dispatched for Neon ARMv7, arch index [3]
in arm_silk_map.c). Replace with vpadd_s32 + lane extraction, which is
portable (compiles on both ARMv7 and AArch64) and produces identical addp
codegen on AArch64 (zero performance impact on the horizontal reduction).

Fixes lpi review:
xiph#483 (comment)
@czoli1976

Copy link
Copy Markdown
Author

Acknowledged the ARMv7 build break you flagged (@lpi): the AArch64-only vaddvq_s32 at the per‑subframe horizontal reduction has been replaced with a portable vpadd_s32 + lane‑extract pattern (commit 1125bbaa). It compiles cleanly on both ARMv7 and AArch64 and produces identical addp codegen on AArch64 (zero perf impact on the reduction, which runs once per 8‑sample block, not in the inner MAC loop).

Re the −0.0578% incremental — that is within run‑to‑run noise here too (the combined #503+#483 result of +4.0% is solid). Will keep an eye on it but no further action needed on the code.

Full analysis of both your review points below:

1. The review, summarized

2. Is the review positive? Yes. lpi confirmed the VAD changes work correctly and the combined improvement is solid +4%. The only actionable items were the ARMv7 guard (fixed) and the noise‑level perf note (not a blocker).

3. The fix
vaddvq_s32 (AArch64 addp) → vpadd_s32 + vget_lane_s32 (portable). See the squash commit for the diff.

@czoli1976

Copy link
Copy Markdown
Author

@lpi do you think it is fine now ? also the ARM v7 is covered

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants