OpenAI has introduced Dots, an always‑on agent that can stay connected to a user’s account without consuming the user’s regular token allowance during the initial launch month. The agent only incurs metered usage when it hands off work to other OpenAI services such as Codex or ChatGPT Work, which continue to count against the usual plan limits.
What Changed?
During the first month after launch, the primary Dot is treated as part of the subscription and does not reduce the user’s allocated usage when it operates autonomously. Once the Dot initiates a task in a separate product—most notably Codex—the operation is billed according to that product’s standard metering. OpenAI’s public terms state that this exemption lasts for the first month, after which per‑plan usage rules will be published. The company also hinted at a future paid tier that would allow higher speed or bandwidth for a Dot, while the baseline remains included.
Why Practitioners Should Care
For AI engineers and platform teams, the distinction between “included” and “metered” work directly affects cost modeling, capacity planning, and budgeting. A Dot can sit idle for hours without impacting the allowance, then trigger charges the moment it delegates to Codex. This behavior introduces a variable cost component that is not visible until the delegation occurs, potentially leading to unexpected consumption spikes.
Architectural and Operational Implications
- Cost predictability: The baseline free‑run time simplifies short‑term budgeting, but any integration that routes work to metered services must be accounted for in cost forecasts.
- Routing decisions: Since a
Dotcan choose between handling work itself, invoking an included tool, or calling a metered product, the chosen path determines whether the user is charged. Engineers need to design explicit routing policies or monitoring to control this behavior. - Monitoring and alerts: The source does not confirm that users receive warnings before a
Dotswitches to a metered operation. Implementing custom telemetry to detect when aDotopens aCodextask can help avoid surprise usage. - Plan changes: OpenAI reduced the included usage on its $200 Pro plan on the same day
Dotslaunched and added a $500 tier. Practitioners must reassess which plan aligns with their expectedDotworkload, especially after the launch‑month exemption ends. - Future paid bandwidth: A forthcoming paid layer for increased speed or bandwidth suggests that high‑throughput or latency‑sensitive workloads may eventually require additional spend beyond the baseline inclusion.
Security and Governance Considerations
Because a Dot can autonomously decide where to execute code, the delegation point becomes a governance boundary. Teams should treat the hand‑off to Codex as a potential attack surface: any compromised Dot could trigger metered operations that consume resources or expose data to downstream services. Auditing the logs of delegation events and enforcing policy checks before a Dot initiates a metered task are prudent safeguards.
Related CloudNinjas coverage: AI engineering.
What This Means For Practitioners
Plan for a two‑phase cost model: treat the always‑on Dot as a free baseline during the launch month, then allocate budget for any delegated calls to Codex or similar services. Implement monitoring that flags the first metered hand‑off per Dot, and consider policy‑driven routing to keep critical workloads on the included path when possible. Finally, stay alert for OpenAI’s upcoming per‑plan usage terms and the announced paid bandwidth tier, as they will reshape the economics of long‑running agents.



