Skip to main content

Perplexity's Hybrid Compute Keeps Sensitive Work Local

Perplexity launched Hybrid Compute for its Mac app, splitting AI tasks between frontier cloud models and local ones running on Apple Silicon. Sensitive work stays on your machine: the app scans uploads, gates anything private to local Gemma and Qwen models, and asks before sending data to the cloud. It is available today to Pro, Max, and enterprise users on macOS 15 or newer.

On this page

What launched today

Perplexity announced Hybrid Compute on Monday as an offshoot of Perplexity Computer, the agent suite the company debuted in February [1]. The idea is simple to describe: a single task can be divided between a frontier cloud model, such as Anthropic's Opus 5 or OpenAI's GPT-5.6 Sol, and a local LLM running on your own machine, with the sensitive portions kept on-device [1].

Availability is immediate but bounded. Hybrid Compute ships to Pro and Max subscribers and enterprise customers, runs only on Apple Silicon Macs, requires macOS 15, and Perplexity recommends at least 32GB of unified memory for the local models [1].

The privacy gate and the local models

The part that will interest privacy-minded users is the gate built into the Mac app. Before delegation begins, Hybrid Compute scans the files and information a task would send, flags anything sensitive, and asks for confirmation before it goes to the cloud. A newly trained privacy classifier suggests which material to keep local [1]. Jon Staff, who oversees Perplexity's Mac products, described it plainly: "we're going to automatically check for sensitive content and make sure that you want to share that data to the cloud" [1].

The local side runs Gemma E4B and two variants of Qwen's 35-billion-parameter 3.6 model, one of them post-trained by Perplexity, installed from inside the app with no terminal required, and more local models are promised [1]. During a task, a visualization shows CPU, GPU, and memory usage while a sidebar tracks token consumption [1].

The cost angle: local tokens are free

Privacy is the headline, but the pricing undercurrent is real. Perplexity states that "you won't be charged for any tokens a local model generates on your personal machine" [1]. For anyone who has watched an agent session burn through a frontier model's quota, routing the grunt work to a local model is a meaningful saving.

Staff was also candid about the trade-off: "a fully frontier output is going to almost always be better in terms of raw artifact creation", while arguing the choice belongs to the user: "I think it's got to be a sliding scale, and we want the user to have control over where they are on that sliding scale" [1]. The launch caps a run of teases: the hybrid local-server orchestrator was first shown at Computex in June [3], and builds of the Hybrid Mode leaked through the summer [2].

What this means for private AI users

This is the first time a mainstream consumer assistant has shipped automatic sensitive-data routing to on-device models, and that normalizes the idea that private does not have to mean separate: the same app, same session, with the private parts handled locally [1]. The classification is Perplexity's own, so treat the gate as a helpful default rather than a guarantee, and review what it proposes to send.

For full control, the options remain what they were: a local runtime such as Ollama or LM Studio keeps every token on your machine, at the cost of doing the wiring yourself [1]. Hybrid Compute sits deliberately in the middle of that sliding scale, and it is a sensible place for most people to start [1].

Sources

  1. Perplexity's Hybrid Compute Splits Sensitive Tasks Between Cloud and Local AIEngadget
  2. Perplexity prepares Hybrid Mode for Computer on MacTestingCatalog
  3. Perplexity AI Introduces Hybrid Local-Server Inference Orchestrator for Personal ComputerMarkTechPost

CuriousLM runs supported AI models locally on your device. Try CuriousLM.