# Evaluating the Applicability of LLM Inference Optimizations on Apple Silicon

**Authors**: Kyrylo Yemets, Andrii Yaroshevych, Victor Muryn, Ostap Khomenko, Pavlo Kryven, Maksym Kmet, Maksym Shamrai, Mariya Hirna  
**Conference**: On-Device Intelligence: Foundation Models under Real-World Constraints (ODI) @ NeurIPS 2026  
**Published**: 2026-09-30  
**Topics**: artificial-intelligence  
**Type**: Paper  
**URL**: https://research.macpaw.com/publications/llm-optimizations-apple-silicon

Apple Silicon is among the most widely used platforms for local LLM inference, yet nearly all inference optimizations are validated on server-class NVIDIA GPUs under batched serving workloads. Which optimizations retain their benefits when moved to Apple Silicon, and what determines their applicability? We evaluate four optimization families: KV-cache pruning, speculative decoding, ported GPU kernels, and post-training quantization. Experiments span four Apple machines under a common single-request decode workload, with matched RTX 4090 comparisons where possible. KV-cache pruning can reach up to 1.75× end-to-end throughput on Apple but can reduce throughput on desktop NVIDIA GPUs. In matched experiments, all evaluated speculative decoding methods reduce throughput on Apple, despite reported 2–6× gains on server GPUs. CUDA kernel designs transfer only partially: public Metal lacks asynchronous copy, and batch-one decode makes several serving-oriented mechanisms inapplicable, although transferable components yield up to 1.20× decode throughput. Quantization gains vary by machine. Across four families, a reported speedup is not a portable property of a technique alone but a joint function of technique, runtime, hardware contract, and workload regime.

```bibtex
@misc{yemets-etal-2026-neurips-odi-apple-silicon,
  author = {Kyrylo Yemets and Andrii Yaroshevych and Victor Muryn and Ostap Khomenko and Pavlo Kryven and Maksym Kmet and Maksym Shamrai and Mariya Hirna},
  title  = {Evaluating the Applicability of {LLM} Inference Optimizations on {Apple Silicon}},
  month  = {September},
  year   = {2026},
  note   = {\emph{Accepted to the On-Device Intelligence Workshop @ NeurIPS 2026.} \url{https://research.macpaw.com/publications/llm-optimizations-apple-silicon}},
}
```
