Live
Treat container images as a security boundary to keep delivery CVE‑freeBackstage AI Integration Takes Center Stage at BackstageCon 2026: Practical Guidance for Platform and Security TeamsAI builder program: Architectural and operational takeaways for engineersClaude Haiku 5.5 slashes token costs and adds effort controls – practical impact for AI workloadsRethinking ROI for Agentic Automation: A Practitioner’s Guide to Value and OperationsOpen‑weight decision models from Cloudflare reshape inference design and opsRedesigning Git Storage for Agent‑Driven Scaling on GitHubCilium networking at AI scale: practical takeaways from CiliumCon 2026Treat container images as a security boundary to keep delivery CVE‑freeBackstage AI Integration Takes Center Stage at BackstageCon 2026: Practical Guidance for Platform and Security TeamsAI builder program: Architectural and operational takeaways for engineersClaude Haiku 5.5 slashes token costs and adds effort controls – practical impact for AI workloadsRethinking ROI for Agentic Automation: A Practitioner’s Guide to Value and OperationsOpen‑weight decision models from Cloudflare reshape inference design and opsRedesigning Git Storage for Agent‑Driven Scaling on GitHubCilium networking at AI scale: practical takeaways from CiliumCon 2026
Anthropic

Claude Haiku 5.5 slashes token costs and adds effort controls – practical impact for AI workloads

AI SummaryPowered by AI

Claude Haiku 5.5 arrives with dramatically lower token prices, effort controls, and benchmark gains. The changes directly affect cost planning, workload orchestration, and security posture for AI‑focused engineering teams.

Anthropic has released Claude Haiku 5.5, the latest iteration of its smallest, cost‑focused model. The update brings a steep token‑price reduction, introduces effort controls, and shows measurable benchmark gains, all of which affect how engineers provision, cost‑track, and secure AI workloads.

Pricing overhaul and cost impact

Haiku 5.5 drops the per‑million‑token rates from the previous $1 input/$5 output to $0.10 input/$0.50 output for requests under 100 k tokens. Larger payloads are billed at $0.50 input/$2.50 output, whereas the older Haiku 4.5 used a flat rate regardless of size. Anthropic notes that roughly 90 % of historic Haiku 4.5 traffic fell into the lower‑priced tier, translating to a 90 % cut for small requests and a 50 % cut for larger ones. The company estimates an average 75 % cost saving when accounting for a mixed workload and a tokenizer that now emits slightly more tokens per task.

Performance improvements

Benchmark data released by Anthropic shows a clear uplift. On the OSWorld 2.1 computer‑use suite, Haiku 5.5 achieved 72.4 % versus 15.7 % for Haiku 4.5 and outperformed GPT‑6 Luna’s 48.9 %. In the GDPval‑AA v2.1 knowledge‑work test, the new model scored 1 620, more than double the 735 points of its predecessor and ahead of Luna’s 1 437. The Terminal‑Bench 4.0 multi‑step command‑line test recorded 39.2 % for Haiku 5.5, a jump from 0 % for Haiku 4.5, though still below Sonnet 5.5’s 70.6 %.

Effort controls and SDK extensions

Haiku 5.5 is the first Haiku variant to expose effort controls, defaulting to a “medium” setting that lets developers cap token consumption per task. The controls are surfaced through the existing Python and TypeScript SDKs, which now include beta support for computer‑use and browser‑use capabilities. This gives teams finer‑grained levers for managing latency and cost in agentic workloads such as live support bots or automated browsing.

Operational and security considerations

The model is available across the Claude Platform, AWS, Google Cloud, and Azure under the identifier claude-haiku-5-5. Deployments must account for the new pricing tiers when estimating budget, especially for batch jobs that exceed 100 k tokens. The tighter cybersecurity safeguards compared with Haiku 4.5 broaden the range of defensive tasks the model can handle, yet penetration‑testing scenarios remain blocked. Organizations requiring broader access can apply to Anthropic’s verification programs, which may affect compliance planning.

Related CloudNinjas coverage: AI engineering.

What This Means For Practitioners

  • Re‑evaluate token‑based cost models; the new rates can reduce spend by up to three‑quarters for typical workloads.
  • Leverage effort controls to enforce predictable latency and token budgets in real‑time agents.
  • Update CI/CD pipelines to reference the claude-haiku-5-5 endpoint and test the beta SDK features before production rollout.
  • Review security policies to incorporate the updated safeguards and consider verification program enrollment if deeper inspection is required.
  • Benchmark against the reported scores if your workloads involve computer use or multi‑step command‑line tasks, to validate the claimed performance uplift.
Originally published atThe New Stack