Live
Integrating Cloudflare’s Web Search API in Beta via AI GatewayOpenTofu migration: practical takeaways ahead of KubeCon North America 2026Claude Code mods enable programmable UI and permission control for AI engineersFine‑tuning Search Agents with Multi‑Turn Reinforcement Learning on SageMaker AIAI Hypercomputer adoption eliminates manual GPU scheduling and slashes job wait times on shared GKE clustersGitHub Copilot code review API adds effort level control and switches default to BalancedGitHub Copilot model deprecation forces updates to AI‑assisted pipelinesProduction‑Ready AI SRE Agents: Architecture, Cost, and Access ShiftsIntegrating Cloudflare’s Web Search API in Beta via AI GatewayOpenTofu migration: practical takeaways ahead of KubeCon North America 2026Claude Code mods enable programmable UI and permission control for AI engineersFine‑tuning Search Agents with Multi‑Turn Reinforcement Learning on SageMaker AIAI Hypercomputer adoption eliminates manual GPU scheduling and slashes job wait times on shared GKE clustersGitHub Copilot code review API adds effort level control and switches default to BalancedGitHub Copilot model deprecation forces updates to AI‑assisted pipelinesProduction‑Ready AI SRE Agents: Architecture, Cost, and Access Shifts
AWS

Fine‑tuning Search Agents with Multi‑Turn Reinforcement Learning on SageMaker AI

AI SummaryPowered by AI

SageMaker AI now offers multi‑turn reinforcement learning for fine‑tuning search agents, enabling smaller models to learn tool‑use across several interaction steps. This gives engineers a cost‑effective path to reliable, multi‑step agent behavior without relying on expensive frontier models.

Amazon SageMaker AI now supports multi‑turn reinforcement learning (MTRL) for fine‑tuning search agents, letting engineers replace costly frontier models with smaller, faster models that learn to use search tools across several interaction steps. This change matters because it offers a way to achieve reliable, multi‑step agent behavior without the latency and expense of large‑scale LLM inference.

Why Multi‑Turn RL Is Different From Prior Approaches

Traditional fine‑tuning either relies on supervised demonstrations, which are expensive to produce, or on single‑turn reinforcement learning that evaluates each response in isolation. Neither method captures the dependencies that arise when an agent must decide which tool to call, interpret results, and decide whether to continue searching. MTRL optimizes the entire sequence of decisions against a final‑outcome reward, aligning training with the actual usage pattern of search agents.

Core Capabilities of SageMaker AI MTRL

  • Modular agent‑environment interface: Engineers define custom reward functions, tool loops, and conversation shapes with minimal code.
  • Serverless execution model: Training runs on a per‑token pricing model without provisioning dedicated GPU clusters.
  • Asynchronous rollout and trajectory collection: Generation and gradient updates occur in parallel, limiting off‑policy staleness while keeping training throughput high.
  • Built‑in algorithm library: Supports Proximal Policy Optimization, Clipped Importance Sampling Policy Optimization, and importance‑sampling losses, together with advantage estimators such as GRPO and RLOO.
  • Resumable jobs: Long training runs can be split across multiple jobs to bypass single‑job time limits.
  • Trajectory observability: MLflow integration lets teams inspect turn‑by‑turn actions and rewards during training.
  • Evaluation jobs: Reward, pass@k, and trajectory metrics can be generated before deploying to an endpoint or Bedrock.

Practical Implementation Sketch

In the example from the blog, a Qwen3.6‑27B model was fine‑tuned for a search agent that can invoke two tools: a BM25 lexical search and a vector‑search service. The environment was set up in the US West (Oregon) region, with datasets stored in S3 and an endpoint exposing the two search tools. Training limited the number of turns to keep agent responses concise and to encourage efficient tool usage.

Typical steps include:

  1. Prepare the rollout environment by implementing a thin wrapper that calls the BM25 and vector search APIs and returns results to the model.
  2. Define a reward function that scores the final answer based on retrieval quality (e.g., relevance or correctness).
  3. Configure a SageMaker MTRL job, selecting an algorithm (e.g., PPO) and any desired advantage estimator.
  4. Launch asynchronous rollouts; the service generates multi‑turn trajectories while a separate worker computes policy gradients.
  5. Monitor trajectory logs in MLflow to verify that the agent is learning to choose the appropriate tool and to stop after an optimal number of turns.
  6. Run an evaluation job to capture pass@k and reward trends before promoting the model to a production endpoint.

Related CloudNinjas coverage: AWS.

What This Means For Practitioners

Engineers can now embed environment‑specific behavior directly into a compact LLM, reducing inference cost and latency while preserving the ability to make complex, multi‑step decisions. Operationally, the serverless, resumable nature of MTRL fits into CI/CD pipelines and avoids long‑running GPU reservations. Observability through MLflow and built‑in evaluation jobs provides a clear feedback loop for model quality and helps surface unexpected tool usage patterns before deployment.

When adopting this workflow, teams should evaluate the trade‑off between model size and the richness of the reward signal, verify that the custom environment wrapper correctly isolates tool calls, and monitor rollout staleness to ensure policy updates stay relevant. Continuous evaluation of pass@k and reward trends will be essential to maintain retrieval quality as the underlying data or search tools evolve.

Originally published atAWS Machine Learning Blog