Key Info
Perplexity has announced that Lily, the local inference engine it built for hybrid compute, is now available. Lily treats Apple silicon as a distinct inference platform and maps Qwen's operations directly to its compute and memory architecture.
Highlights
- Lily is specialized for running Qwen3.6-35B-A3B on Apple silicon, using the chip's architecture as the basis for inference.
- Designed for on-device compute in Perplexity Computer, Lily is now publicly accessible via the link in the announcement.