Perplexity Open Sources Lily: A Rust + Metal Inference Engine for Qwen3.6-35B-A3B on Apple Silicon
Perplexity AI has released Lily, an open-source inference engine optimized for running large language models on Apple Silicon devices. Built using Rust and custom Metal kernels, Lily powers the Hybrid Compute feature in Perplexity Computer and delivers significant performance improvements over existing solutions. The engine is specifically optimized for the Qwen 3.6-35B-A3B model on Apple's latest processors, representing a focused approach to specialized hardware optimization.
Lily demonstrates substantial throughput advantages compared to MLX-LM, the current industry benchmark for Apple Silicon inference. On a 40-core, 128 GB M5 Max processor, Lily achieves 1.23x faster prefill throughput and 1.35x faster decode throughput. These metrics measure how quickly the engine processes initial input tokens and generates subsequent responses, respectively. The specialized implementation leverages Rust's performance characteristics alongside Metal, Apple's native graphics programming framework, enabling efficient utilization of Apple Silicon's neural processing capabilities. By targeting a specific model-processor combination, Perplexity prioritized optimization depth over broad compatibility.
- Local Inference Advancement: Lily accelerates the trend toward privacy-preserving, on-device AI inference without cloud dependencies
- Open-Source Contribution: The release strengthens the Apple Silicon AI ecosystem, benefiting developers building consumer and professional applications
- Performance Standards: Establishes new efficiency baselines for metal-optimized inference, potentially influencing future tooling development
- Hardware Specialization: Demonstrates commercial viability of purpose-built inference engines targeting specific processor architectures
- Competitive Differentiation: Provides Perplexity's products with measurable technical advantages in local computation speed
The open-sourcing of Lily marks an important moment for edge AI development, particularly as consumer devices become more powerful. By releasing specialized inference technology, Perplexity validates the market demand for efficient local language model execution while contributing infrastructure that the broader developer community can study and build upon. This move supports the growing momentum toward decentralized AI computing, where users maintain greater privacy and control while reducing latency and cloud service dependencies. As large language models continue evolving, purpose-built inference engines like Lily will likely become essential for delivering responsive, private AI experiences on consumer hardware.
Key Takeaways
- Perplexity AI has released Lily, an open-source inference engine optimized for running large language models on Apple Silicon devices.
- Built using Rust and custom Metal kernels, Lily powers the Hybrid Compute feature in Perplexity Computer and delivers significant performance improvements over existing solutions.
- The engine is specifically optimized for the Qwen 3.
- 6-35B-A3B model on Apple's latest processors, representing a focused approach to specialized hardware optimization.
Read the full article on MarkTechPost
Read on MarkTechPost