Live
OpenTelemetry tracing expands across Cloudflare’s proxy stack in betaDynamic Model Triage: Engineering Implications of Grok Bot’s Multi‑Model BackendAccess Cloudflare Skills Directly Through the API MCP ServerCodeQL 2.27.2 expands language models and tightens macOS build support – what engineers need to knowTangible Certification: Turning a Kubernetes Badge into a Gold NecklaceGoogle Data Cloud GA updates: agent‑centric tooling, hybrid Spanner, and expanded Lakehouse catalogCloudflare folds Deno runtime into Workers: practical impact on serverless deploymentsManaging Copilot Code Review Costs and License Scope with New Org‑Level ControlsOpenTelemetry tracing expands across Cloudflare’s proxy stack in betaDynamic Model Triage: Engineering Implications of Grok Bot’s Multi‑Model BackendAccess Cloudflare Skills Directly Through the API MCP ServerCodeQL 2.27.2 expands language models and tightens macOS build support – what engineers need to knowTangible Certification: Turning a Kubernetes Badge into a Gold NecklaceGoogle Data Cloud GA updates: agent‑centric tooling, hybrid Spanner, and expanded Lakehouse catalogCloudflare folds Deno runtime into Workers: practical impact on serverless deploymentsManaging Copilot Code Review Costs and License Scope with New Org‑Level Controls

Portable Computer brings Perplexity’s local AI agent to the desktop – what engineers need to know

AI SummaryPowered by AI

Perplexity’s Portable Computer moves its agentic AI assistant from the cloud to local Nvidia‑based hardware, introducing a new harness, deterministic orchestrator, and OS‑level sandbox. This change forces engineers to reassess GPU capacity, sandbox constraints, and cloud‑fallback approval processes when integrating powerful AI agents into on‑premises workflows.

Perplexity has released Portable Computer, a desktop‑ready version of its agentic AI assistant that runs entirely on local hardware. The shift from a cloud‑only service to a self‑hosted model changes the compute, security, and operational landscape for engineers who need to integrate powerful AI tools without sending data off‑premises.

Hardware and runtime requirements

The solution is limited to two hardware paths. The first is Nvidia’s DGX Spark workstation, priced around $4,800, which ships with a specialised OS. The second path accepts a standard Ubuntu machine—ARM or x64—equipped with an Nvidia RTX GPU that provides at least 24 GB of VRAM; a current RTX 3090 with that capacity costs roughly $1,500. No other GPU class is supported, and the product is not yet available for macOS or Windows, though Windows support is slated for a later release.

Architecture: harness, orchestrator, and sandbox

Running the agent locally required a redesign of the software stack. Perplexity kept most of the agent’s capabilities but introduced a new harness that adapts model configuration to the host’s resources. The harness delegates task planning to the Qwen3.8-27B model while a deterministic orchestrator—implemented as plain code rather than another model—assembles context, enforces policy, and executes tool calls inside an OS‑level sandbox. The sandbox limits process creation, file‑system paths, and network egress. If the sandbox cannot be instantiated, the harness disables tool execution entirely, preventing uncontrolled system access.

Performance and fallback model interaction

Perplexity reports that the local model, with a 260 k token context window but practical limits near 100 k tokens, achieves higher scores on internal benchmarks than comparable agents. On the 53‑task Local Knowledge Work Bench, Portable Computer reached 82.6 % accuracy, outperforming Pi (77.6 %) and Hermes (74 %). On the ParseBench‑100 suite, it scored 65.1 % versus Hermes at 34.6 % and Pi at 13.9 %.

When a task exceeds the local model’s capability, the harness can invoke a cloud‑based model for guidance. Before any data leaves the machine, the harness extracts the relevant context, flags potentially sensitive content, and prompts the user for approval. The cloud model returns textual advice only; it never receives direct file or tool access, and the local orchestrator incorporates the advice back into the ongoing run.

Connectors for Google Drive, Gmail, Slack, and GitHub are bundled, allowing the agent to interact with external services while keeping inference and private‑document processing on the host. Web searches and connector calls are the only operations that exit the device.

Related CloudNinjas coverage: AI engineering.

What This Means For Practitioners

Engineers should evaluate whether existing GPU resources meet the 24 GB VRAM threshold and consider the cost of a DGX Spark if a dedicated appliance is required. The sandbox model mandates that any custom tool integration respect the same process, file, and network constraints, which may affect existing CI/CD pipelines or automation scripts. The deterministic orchestrator provides a clear audit point for policy enforcement, but teams must still manage the approval workflow for cloud fallback calls to avoid inadvertent data leakage. Finally, the current platform support limits adoption to Linux environments; planning for Windows rollout will be necessary for mixed‑OS fleets.

Originally published atThe New Stack