Perplexity Releases Lily to Speed Up Qwen3.6 on Mac
Perplexity has open-sourced Lily, a highly specialized local inference engine that runs the Qwen3.6-35B-A3B model on Apple Silicon up to 1.35 times faster than Apple's own MLX-LM framework.

Perplexity has launched Lily, an open-source local inference engine written in Rust with custom Metal kernels. Designed specifically for the Qwen3.6-35B-A3B model on Apple Silicon, Lily bypasses PyTorch and MLX entirely to power the local portion of Perplexity Computer's hybrid setup. The 19.4 GB checkpoint uses groupwise 4-bit weights and features a complex architecture combining 10 full-attention layers with 30 Gated DeltaNet layers, alongside a sparse mixture-of-experts router that selects eight of 256 experts plus one shared expert per token.
Tested on an M5 Max MacBook Pro with a 40-core GPU and 128 GB of unified memory, Lily achieved significant performance gains over Apple's MLX-LM. It averaged 1.23 times the prefill throughput and 1.35 times the decode throughput across context lengths from 256 to 128K tokens. At a 4K prompt and 4K decode context, Lily reached 5,749.9 prefill tokens per second and 186.6 decode tokens per second, compared to MLX-LM's 4,737.5 and 140.9. This speedup comes with virtually no loss in quality; Lily's perplexity was only 0.04 percent higher, matching the top-ranked token in 96.35 percent of positions.
To achieve these speeds, the engine implements several low-level optimizations. Fusing 4-bit dequantization into the grouped GEMM boosted prefill throughput by 77.4 percent at a 512-token prompt, while keeping the mixture-of-experts routing on the GPU added another 89 percent. For decoding, Grouped Query Attention packing yielded a 23.8 percent improvement at a 32K context, and switching attention layouts at longer contexts provided gains of 7.7 percent at 32K, 27.4 percent at 64K, and 40.2 percent at 128K. However, speculative decoding proved counterproductive, slowing batch-1 decode by 18 percent.
For practitioners, Lily demonstrates that extreme hardware-model co-design can push local consumer hardware to its absolute limits, with its kernels reaching up to 97.9 percent of maximum memory bandwidth. While this highly specialized engine loses its advantages if ported to other models, it provides a powerful blueprint for running complex, multi-turn agentic workloads locally on Mac hardware without relying on cloud APIs.
This is our own summary of reporting by AlphaSignal



