Pi has moved the Model Context Protocol (MCP) from an optional add‑on into the core of its coding agent, but it now hides the bulk of tool definitions behind a lightweight sandbox called Codemode. This change slashes the number of tokens consumed before any useful work begins, which directly impacts the usable context window for large language model (LLM) calls.
What Changed in Pi’s MCP Integration
Earlier versions of Pi left MCP servers fully exposed to the model. A single server, such as Chrome DevTools MCP, required roughly 18,000 tokens just to describe its 21 tools, eating about 9 % of a 200,000‑token context window. Playwright MCP used a similar amount (≈13,700 tokens). The model also had to re‑ingest every tool’s result to persist or combine data, adding further overhead.
Pi 1.0 now embeds MCP but, by default, only inserts a one‑line summary of each server into the system prompt. The actual tool definitions live in Codemode, a QuickJS sandbox that runs without Node APIs, file‑system, network, or timers. Scripts inside Codemode can discover, invoke, and post‑process tools before returning a concise result to the model. Developers can still override this behavior with a toolExposure setting to expose, hide, or block individual tools.
Why Token Overhead Matters for AI‑Driven Tooling
LLM providers charge per token and impose hard limits on context size. Consuming 18 k tokens on static tool metadata reduces the space available for user prompts, code, and intermediate results, leading to higher costs and more frequent truncation. By moving most declarations into a 3,000‑token budget that stays discoverable but out of the prompt, Pi cuts the default prompt size from about 5,300 tokens to 3,300 tokens for a GPT‑5.6 request. The savings are especially relevant for agents that chain many tools or operate in long‑running sessions.
Architecture and Operational Implications
Codemode introduces a clear separation between the LLM and the tool execution environment. The sandboxed JavaScript runtime prevents accidental exposure of host resources, which can simplify security reviews. However, practitioners must manage the 3,000‑token declaration budget: tools beyond that limit remain discoverable but are not pre‑declared, which could affect latency for first‑use calls.
Exposure controls let teams tailor the surface area per service. For example, a GitHub integration might expose search_code directly, keep get_* behind Codemode for post‑processing, and block delete_* entirely. This granularity supports least‑privilege principles without altering the underlying MCP server.
Operationally, the reduced prompt size means fewer retries due to context overflow and lower token‑based billing. Teams should monitor token consumption per request and adjust toolExposure or refactor large toolsets into smaller, composable scripts when the budget is approached.
Related CloudNinjas coverage: AI engineering.
What This Means For Practitioners
Adopting Pi’s Codemode‑backed MCP approach can reclaim a significant portion of the LLM context window, lower costs, and provide finer control over tool exposure. Engineers should evaluate existing MCP servers for token bloat, consider moving heavyweight tool definitions into sandboxed scripts, and configure toolExposure to match security and performance goals. Ongoing monitoring of token budgets and prompt sizes will be essential to maintain efficient, composable AI workflows.

