Conversation
|
Hi, @jmvalin Could you please help review this PR when you have a moment? It includes targeted NSQ_del_dec_neon_intr.c Multiple optimizations,for details, see the commit message. Please let me know if you need any additional information or if there are any changes required. Thanks in advance for your time! |
|
Tested head: I tested safe ideas independently on top of #503. Their approximate combined production-c10 encoder improvements versus stock were:
Positive means faster/lower per-call cost. These stock-relative numbers are compounded cross-run estimates, not direct stock arms. Only LF_AR caching added a small increment over #503: +0.0768% CPU / +0.0616% wall. I did not time the submitted bundle wholesale because its inline assembly does not model pointed-to memory, its prefetch pointer formation extends beyond the derived state-row region, and its state-copy rewrite overlaps #503. The allpass-unroll and AR-tail ideas compiled to equivalent Release code. The tested isolated variants were output-exact and sanitizer-clean. |
silk/arm/NSQ_del_dec_neon_intr.c — Multiple optimizations