GitHub Copilot CLI version 1.0.94-0 now includes a local model discovery workflow that lets you query a running Ollama instance for compatible models directly from the CLI. This change removes the need to leave your terminal or pre‑configure models before you can experiment with a locally hosted LLM.
What Changed
The /model command now lists models served by Ollama alongside any cloud‑based models you have configured. Discovery is read‑only – it does not add the model automatically. After selecting a model you can either add it for the current session and start using it immediately, or add it without switching the active model. The CLI picks up the new model without a restart.
Why Practitioners Should Care
AI engineers and platform teams can iterate on locally built or fine‑tuned models without disrupting existing workflows. DevOps and SRE staff gain a faster feedback loop for testing model behavior in CI pipelines or on‑prem environments. Security engineers see that opting for a local model does not implicitly enable offline mode or suppress telemetry, preserving visibility into usage while still allowing isolated execution when desired.
Architectural and Operational Implications
- Dependency Management: Ollama and the target model must already be installed; the CLI does not perform installation or download actions.
- Model Capabilities: Only models that support tool calling and streaming appear in the picker, limiting the set to those that can handle Copilot’s interactive features.
- Provider Errors: Connection failures to a local provider surface in the selection UI with an explanatory message, giving operators immediate diagnostic context.
- Offline Mode Separation: Enabling
COPILOT_OFFLINE=trueremains an explicit step. Selecting a local model does not toggle offline mode, and remote providers can still receive prompts over the network even when offline mode is set. - Telemetry Continuity: GitHub telemetry remains active unless explicitly disabled, ensuring that usage data continues to flow for monitoring and compliance.
Related CloudNinjas coverage: AI engineering.
What This Means For Practitioners
Adopt the new /model workflow to evaluate locally hosted LLMs without altering your CI/CD pipelines or requiring a CLI restart. Verify that your Ollama deployment satisfies the tool‑calling and streaming requirements before adding it to the picker. Keep offline mode and telemetry settings explicit in your environment configuration to avoid accidental exposure or loss of observability. Finally, monitor the upcoming intelligent routing feature announced by Microsoft for potential automation of model selection once it becomes generally available.


