Perplexity has released Portable Computer, a desktop‑ready version of its agentic AI assistant that runs entirely on local hardware. The shift from a cloud‑only service to a self‑hosted model changes the compute, security, and operational landscape for engineers who need to integrate powerful AI tools without sending data off‑premises.
Hardware and runtime requirements
The solution is limited to two hardware paths. The first is Nvidia’s DGX Spark workstation, priced around $4,800, which ships with a specialised OS. The second path accepts a standard Ubuntu machine—ARM or x64—equipped with an Nvidia RTX GPU that provides at least 24 GB of VRAM; a current RTX 3090 with that capacity costs roughly $1,500. No other GPU class is supported, and the product is not yet available for macOS or Windows, though Windows support is slated for a later release.
Architecture: harness, orchestrator, and sandbox
Running the agent locally required a redesign of the software stack. Perplexity kept most of the agent’s capabilities but introduced a new harness that adapts model configuration to the host’s resources. The harness delegates task planning to the Qwen3.8-27B model while a deterministic orchestrator—implemented as plain code rather than another model—assembles context, enforces policy, and executes tool calls inside an OS‑level sandbox. The sandbox limits process creation, file‑system paths, and network egress. If the sandbox cannot be instantiated, the harness disables tool execution entirely, preventing uncontrolled system access.
Performance and fallback model interaction
Perplexity reports that the local model, with a 260 k token context window but practical limits near 100 k tokens, achieves higher scores on internal benchmarks than comparable agents. On the 53‑task Local Knowledge Work Bench, Portable Computer reached 82.6 % accuracy, outperforming Pi (77.6 %) and Hermes (74 %). On the ParseBench‑100 suite, it scored 65.1 % versus Hermes at 34.6 % and Pi at 13.9 %.
When a task exceeds the local model’s capability, the harness can invoke a cloud‑based model for guidance. Before any data leaves the machine, the harness extracts the relevant context, flags potentially sensitive content, and prompts the user for approval. The cloud model returns textual advice only; it never receives direct file or tool access, and the local orchestrator incorporates the advice back into the ongoing run.
Connectors for Google Drive, Gmail, Slack, and GitHub are bundled, allowing the agent to interact with external services while keeping inference and private‑document processing on the host. Web searches and connector calls are the only operations that exit the device.
Related CloudNinjas coverage: AI engineering.
What This Means For Practitioners
Engineers should evaluate whether existing GPU resources meet the 24 GB VRAM threshold and consider the cost of a DGX Spark if a dedicated appliance is required. The sandbox model mandates that any custom tool integration respect the same process, file, and network constraints, which may affect existing CI/CD pipelines or automation scripts. The deterministic orchestrator provides a clear audit point for policy enforcement, but teams must still manage the approval workflow for cloud fallback calls to avoid inadvertent data leakage. Finally, the current platform support limits adoption to Linux environments; planning for Windows rollout will be necessary for mixed‑OS fleets.


