Skip to content

Enabling neon intrinsics for silk_NSQ_del_dec_c() on Apple aarch64 for 34% speedup on M2 - #501

Open
knutinh wants to merge 1 commit into
xiph:mainfrom
knutinh:patch-1
Open

knutinh wants to merge 1 commit into
xiph:mainfrom
knutinh:patch-1

Conversation

@knutinh

@knutinh knutinh commented Sep 14, 2026

Copy link
Copy Markdown

Motivation:
While profiling Opus encode on an ARM-equipped Apple laptop, I noticed that the most expensive function (by about 70%) was silk_NSQ_del_dec_c() and that the codebase contains a neon version of the same function, silk_NSQ_del_dec_neon(). I was curious as to why that was not used, and how that affected performance.

Inband FEC doubles the calls to NSQ per frame, thus exaggerating the performance diff compared to a non-FEC encode.

Source input: official testvectors, decoded, concatenated and repeated
Encoder settings: complexity=9, FEC on (30% loss), silk, bitrate 64k
System:MacBook Pro 2023 Apple M2 Pro, Tahoe 26.5.2

** What I did: **
% Fetch and build baseline:
git clone https://gitlab.xiph.org/xiph/opus.git
./autogen.sh (for getting DNN models)

meson setup builddir-baseline
ninja -C builddir-baseline
meson test -C builddir-baseline

%download testvectors, concatenate and repeat for length
curl -O https://opus-codec.org/static/testvectors/opus_testvectors.tar.gz
tar xf opus_testvectors.tar.gz
for i in 01 02 03 04 05 06 07 08 09 10 11 12; do
./builddir-baseline/src/opus_demo -d 48000 1 opus_testvectors/testvector${i}.bit tv${i}.pcm
done
cat tv*.pcm > base.pcm
for i in $(seq 10); do cat base.pcm; done > speech_long.pcm

% Test performance:
time ./builddir-baseline/src/opus_demo -e restricted-silk 48000 1 64000 -inbandfec -loss 30 -complexity 9 speech_long.pcm /dev/null
34.55s user 0.30s system 99% cpu 34.925 total

% apply patch:

meson setup builddir-proposed
ninja -C builddir-proposed
meson test -C builddir-proposed

% Verify bit-exact output:
./builddir-baseline/src/opus_demo -e restricted-silk 48000 1 64000 -inbandfec -loss 30 -complexity 9 speech_long.pcm /tmp/base.bit
./builddir-proposed/src/opus_demo -e restricted-silk 48000 1 64000 -inbandfec -loss 30 -complexity 9 speech_long.pcm /tmp/prop.bit
cmp /tmp/base.bit /tmp/prop.bit && echo "bit-exact"
bit-exact

% Verify binary diff
cmp builddir-baseline/src/libopus.0.dylib builddir-proposed/src/libopus.0.dylib
builddir-baseline/src/libopus.0.dylib builddir-proposed/src/libopus.0.dylib differ: char 1001, line 1

% Test new performance:
time ./builddir-proposed/src/opus_demo -e restricted-silk 48000 1 64000 -inbandfec -loss 30 -complexity 9 speech_long.pcm /dev/null
22.97s user 0.35s system 99% cpu 23.341 total

% speedup:
(1-22.97/34.55)*100 =34%

% sampling
sample 10 -f /tmp/prof_baseline.txt
1002772/8463=33% samples in silk_NSQ_del_dec_c() from main-frame encoding
100
3341/8463=39% samples in silk_NSQ_del_dec_c() from LBRR-frame encoding

sample 10 -f /tmp/prof_proposed.txt
1002147/8501 =25% samples in silk_NSQ_del_dec_neon() from main-frame encoding
100
2648/8501 =31% samples in silk_NSQ_del_dec_neon() from LBRR-frame encoding

About the proposed patch:
All other ARM NEON dispatch headers in the tree guard this branch with
both conditions:

silk/arm/biquad_alt_arm.h:43 #if !defined(OPUS_HAVE_RTCD) && defined(OPUS_ARM_PRESUME_NEON)
silk/arm/LPC_inv_pred_gain_arm.h:39 (same)
celt/arm/pitch_arm.h:38 (same)

silk/arm/NSQ_del_dec_arm.h:47 only tests !OPUS_HAVE_RTCD. On AArch64
builds without RTCD, PRESUME_NEON() expands via PRESUME_MEDIA() to the
_c variant (celt/arm/armcpu.h:68-70), so OVERRIDE_silk_NSQ_del_dec gets
defined and blocks the correct OPUS_ARM_PRESUME_NEON_INTR branch below.
The patch aligns this file with the other three.

Enabling neon intrinsics on Apple aarch64
@jmvalin

jmvalin commented Sep 15, 2026

Copy link
Copy Markdown
Member

Isn't the fundamental problem that OPUS_ARM_PRESUME_NEON isn't defined? AArch64 guarantees Neon, so presumably that should be set. If it's not then it may be causing other (less obvious) issues.

@lpi

lpi commented Sep 16, 2026 •

Copy link
Copy Markdown

Tested head: 240af26654d8289536a457d640c0101401d07a70.

Under the current CMake AArch64 configuration, the incremental expectation is 0%: baseline and this PR already select the existing _neon implementation, and their normalized NSQ codegen is identical.

In the tested affected default-Autotools and scoped-Meson configurations, the change instead switches the call from _c to the existing _neon implementation. That scalar-to-NEON configuration was not timed, so no honest stock-relative percentage is available.

This is build/dispatch enablement rather than a new optimized kernel. A broader AArch64 PRESUME_NEON configuration fix would be preferable to treating the local guard change as the complete solution.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants