The latest Ollama release (v0.33.0) adds a dedicated local proxy that restores the Claude Desktop Ollama integration, allowing the Claude Desktop client to forward prompts to any model listed in Ollama’s catalog – for example Qwen, DeepSeek and Kimi – whether the model runs locally or in Ollama Cloud.
How the integration works
Ollama now exposes a toggle labeled “Use Ollama models” in its macOS menu. When enabled, the Ollama app automatically configures Claude Desktop’s third‑party gateway to point at the local proxy. The proxy translates Claude’s request format into Ollama’s API, so the model picker inside Claude Desktop shows every model that Ollama can serve. Users can map Claude’s named slots (e.g., “Opus 5” or “Sonnet 5”) to specific underlying models such as Kimi K3 or DeepSeek V4 Pro. Turning the toggle off restores the original Anthropic‑only configuration.
Operational considerations
- Switching between Anthropic and Ollama models is a UI action; no manual endpoint reconfiguration is required.
- Because the proxy runs locally, latency is limited to the host machine’s compute capacity and any network hop to Ollama Cloud if a remote model is selected.
- Model selection can be driven by cost, speed, or the need for fine‑tuned variants that developers have trained on their own data.
- The feature is currently limited to the macOS Ollama client; Windows support has been hinted at but is not yet available.
Security and data‑flow implications
Routing Claude Desktop traffic through a local proxy means that prompt data can stay on the developer’s machine when a locally‑hosted model is chosen, reducing exposure to external services. Conversely, selecting an Ollama Cloud model re‑introduces outbound traffic to Ollama’s hosted endpoints, so teams must evaluate the trust model of that service. The integration does not alter Claude Desktop’s built‑in “Auto mode” permission flow, so any user‑prompted consent mechanisms remain unchanged.
Related CloudNinjas coverage: AI engineering.
What This Means For Practitioners
Engineers can now experiment with a broader set of LLMs without abandoning the Claude Desktop workflow, giving them flexibility to optimise for cost, performance, or custom data. Adoption should be accompanied by a review of where model inference occurs (local vs. cloud), the associated resource requirements, and the data‑privacy posture of any remote Ollama endpoints. Teams should also monitor the upcoming Windows support if cross‑platform consistency is required.

