Conversation
Enabling neon intrinsics on Apple aarch64
|
Isn't the fundamental problem that OPUS_ARM_PRESUME_NEON isn't defined? AArch64 guarantees Neon, so presumably that should be set. If it's not then it may be causing other (less obvious) issues. |
|
Tested head: Under the current CMake AArch64 configuration, the incremental expectation is 0%: baseline and this PR already select the existing In the tested affected default-Autotools and scoped-Meson configurations, the change instead switches the call from This is build/dispatch enablement rather than a new optimized kernel. A broader AArch64 |
Motivation:
While profiling Opus encode on an ARM-equipped Apple laptop, I noticed that the most expensive function (by about 70%) was silk_NSQ_del_dec_c() and that the codebase contains a neon version of the same function, silk_NSQ_del_dec_neon(). I was curious as to why that was not used, and how that affected performance.
Inband FEC doubles the calls to NSQ per frame, thus exaggerating the performance diff compared to a non-FEC encode.
Source input: official testvectors, decoded, concatenated and repeated
Encoder settings: complexity=9, FEC on (30% loss), silk, bitrate 64k
System:MacBook Pro 2023 Apple M2 Pro, Tahoe 26.5.2
** What I did: **
% Fetch and build baseline:
git clone https://gitlab.xiph.org/xiph/opus.git
./autogen.sh (for getting DNN models)
meson setup builddir-baseline
ninja -C builddir-baseline
meson test -C builddir-baseline
%download testvectors, concatenate and repeat for length
curl -O https://opus-codec.org/static/testvectors/opus_testvectors.tar.gz
tar xf opus_testvectors.tar.gz
for i in 01 02 03 04 05 06 07 08 09 10 11 12; do
./builddir-baseline/src/opus_demo -d 48000 1 opus_testvectors/testvector${i}.bit tv${i}.pcm
done
cat tv*.pcm > base.pcm
for i in $(seq 10); do cat base.pcm; done > speech_long.pcm
% Test performance:
time ./builddir-baseline/src/opus_demo -e restricted-silk 48000 1 64000 -inbandfec -loss 30 -complexity 9 speech_long.pcm /dev/null
34.55s user 0.30s system 99% cpu 34.925 total
% apply patch:
meson setup builddir-proposed
ninja -C builddir-proposed
meson test -C builddir-proposed
% Verify bit-exact output:
./builddir-baseline/src/opus_demo -e restricted-silk 48000 1 64000 -inbandfec -loss 30 -complexity 9 speech_long.pcm /tmp/base.bit
./builddir-proposed/src/opus_demo -e restricted-silk 48000 1 64000 -inbandfec -loss 30 -complexity 9 speech_long.pcm /tmp/prop.bit
cmp /tmp/base.bit /tmp/prop.bit && echo "bit-exact"
bit-exact
% Verify binary diff
cmp builddir-baseline/src/libopus.0.dylib builddir-proposed/src/libopus.0.dylib
builddir-baseline/src/libopus.0.dylib builddir-proposed/src/libopus.0.dylib differ: char 1001, line 1
% Test new performance:
time ./builddir-proposed/src/opus_demo -e restricted-silk 48000 1 64000 -inbandfec -loss 30 -complexity 9 speech_long.pcm /dev/null
22.97s user 0.35s system 99% cpu 23.341 total
% speedup:
(1-22.97/34.55)*100 =34%
% sampling
sample 10 -f /tmp/prof_baseline.txt
1002772/8463=33% samples in silk_NSQ_del_dec_c() from main-frame encoding
1003341/8463=39% samples in silk_NSQ_del_dec_c() from LBRR-frame encoding
sample 10 -f /tmp/prof_proposed.txt
1002147/8501 =25% samples in silk_NSQ_del_dec_neon() from main-frame encoding
1002648/8501 =31% samples in silk_NSQ_del_dec_neon() from LBRR-frame encoding
About the proposed patch:
All other ARM NEON dispatch headers in the tree guard this branch with
both conditions:
silk/arm/biquad_alt_arm.h:43 #if !defined(OPUS_HAVE_RTCD) && defined(OPUS_ARM_PRESUME_NEON)
silk/arm/LPC_inv_pred_gain_arm.h:39 (same)
celt/arm/pitch_arm.h:38 (same)
silk/arm/NSQ_del_dec_arm.h:47 only tests !OPUS_HAVE_RTCD. On AArch64
builds without RTCD, PRESUME_NEON() expands via PRESUME_MEDIA() to the
_c variant (celt/arm/armcpu.h:68-70), so OVERRIDE_silk_NSQ_del_dec gets
defined and blocks the correct OPUS_ARM_PRESUME_NEON_INTR branch below.
The patch aligns this file with the other three.