AWS has documented a SageMaker AI multi-turn reinforcement learning (MTRL) workflow for fine-tuning a Qwen3.6-27B search agent, reporting improved retrieval scores on three of four held-out benchmarks and substantially fewer failures in its configured environment. For teams building tool-using search agents, the guide shows how a final retrieval-quality reward and explicit budget-failure penalty can shape behavior across an entire search sequence rather than optimizing each response independently.

The post is an implementation report, not a new model release. AWS positions MTRL as a way to teach a smaller model the tools and interaction patterns of a particular environment, potentially avoiding the latency and cost associated with using a frontier model for every search task. The reported measurements are AWS’s own evaluation runs, however, and apply to the datasets, tools and limits used in this example.

How the multi-turn training design works

Amazon SageMaker AI MTRL treats an agentic task as a series of decisions. It uses multi-turn rollouts to generate training data and policy-gradient methods to optimize the model over the resulting trajectory. Rather than requiring expert demonstrations of ideal search paths, the approach can reward the final outcome after the agent has made several search decisions.

AWS says the service lets users define custom rewards, tool loops and conversation shapes. It also offers serverless execution priced per token, asynchronous rollout and trajectory collection, a choice of policy-optimization and advantage-estimation methods, resumable jobs, MLflow-managed trajectory and reward inspection, and evaluation jobs before deployment to SageMaker AI endpoints or Amazon Bedrock.

In the documented configuration, the search agent could choose between BM25 lexical search and vector search. BM25 was used for exact terms or identifiers, while vector search was intended for semantic or conceptual queries. The setup capped the number of interaction turns to discourage unnecessarily long searches; it required an agent endpoint exposing those tools, training and validation data in Amazon S3, and SageMaker AI access in US West (Oregon).

Reward, data and job configuration

AWS trained on FRAMES, BRIGHT, Enterprise RAG, ESCI, Musique and MLQA, reserving 5% of each training dataset for validation. It evaluated the fine-tuned model on FreshStack, WixQA, BrowseComp-Plus and Wands. Those test sets span developer-oriented questions, help-center support questions, deep-research queries and product-search relevance judgments.

The primary reward was nDCG@10, a ranking metric that measures how closely the top 10 retrieved documents match an ideal relevance ordering. AWS applied it to the completed multi-turn trajectory, so the reward reflected the final retrieved documents rather than an isolated tool call. It also assigned a reward of -1 when the agent exhausted its turn limit or maximum sampling-token allowance in a turn, making successful completion within the configured budget part of the training target.

AWS says it changed only three job settings: one training epoch, a global batch size of 128 prompts and rollout concurrency of 32. Algorithm selection, the advantage estimator and off-policy-staleness bounds remained at defaults. The reported training and validation nDCG@10 curves rose and then plateaued. AWS notes that MTRL jobs have a default 24-hour limit, adjustable through the CreateJob schema, and can resume from checkpoints after timeouts or infrastructure errors.

Reported test results

Against the base Qwen3.6-27B, the fine-tuned agent’s nDCG@10 rose from 0.5725 to 0.6781 on the 400-question WixQA benchmark, from 0.5762 to 0.6112 on the 147-question Wands benchmark, and from 0.5136 to 0.6354 on the 830-question BrowseComp-Plus benchmark. AWS describes those changes as gains of 18.4%, 6% and 23.7%, respectively.

The largest reliability change appeared on BrowseComp-Plus. Its failure rate declined from 22.89% for the base model to 0.68% after fine-tuning, while average turns fell from 7.0 to 6.3. On WixQA, failures declined from 0.67% to 0.17%, although average turns increased from 4.3 to 4.5. Wands recorded a 0% failure rate for both versions, with average turns increasing from 2.2 to 2.9.

The outcome was not uniform across all tests. On FreshStack, nDCG@10 slipped from 0.4112 to 0.4089 after fine-tuning, even though the failure rate fell from 0.20% to 0.05% and average turns decreased from 3.1 to 2.8. Because failed tasks received an nDCG@10 score of zero in the evaluation, the retrieval metric incorporates both ranking performance and task-completion failures.

What the guide suggests for practitioners

The practical lesson in AWS’s example is to make the final task metric drive optimization across the whole agent trajectory, then separately penalize the operational failures that matter in production. For a search agent, that means defining a retrieval-quality measure, exposing the relevant retrieval tools through an endpoint, preparing data in the required MTRL format and evaluating the fine-tuned result against a base model on held-out tasks.

AWS recommends starting with default configurations, defining a task-specific reward and iterating after evaluation. The case study supports that procedure for the reported environment, but it does not establish that every search-agent deployment will see the same retrieval or reliability gains; FreshStack’s small regression is a direct reminder that results can vary by benchmark.

Source: AWS Machine Learning Blog

Definition. Multi-turn reinforcement learning optimizes an agent across a sequence of tool-use decisions using rewards tied to the final task outcome.

BenchmarkReported nDCG@10, base to fine-tuned
WixQA0.5725 to 0.6781
Wands0.5762 to 0.6112
BrowseComp-Plus0.5136 to 0.6354
FreshStack0.4112 to 0.4089

Key takeaways

  • AWS applied final-trajectory nDCG@10 as the primary reward and imposed a -1 reward for exhausting configured turn or token budgets.
  • The fine-tuned agent improved nDCG@10 on WixQA, Wands and BrowseComp-Plus, but slipped slightly on FreshStack.
  • BrowseComp-Plus recorded the largest reliability change, with failures falling from 22.89% to 0.68%.
  • The example used BM25 for exact terms or identifiers and vector search for semantic or conceptual queries.
  • AWS says results apply to the datasets, tools and limits in its example and may not generalize to every search-agent deployment.

FAQ

What did AWS fine-tune with multi-turn RL?

AWS documented a SageMaker AI workflow for fine-tuning a Qwen3.6-27B search agent.

What reward did the AWS workflow use?

Its primary reward was nDCG@10 applied to the completed multi-turn trajectory, with a -1 reward for exhausting the configured turn limit or maximum sampling-token allowance in a turn.

Which benchmarks improved after fine-tuning?

The reported nDCG@10 improved on WixQA, Wands and BrowseComp-Plus, while FreshStack declined slightly.

Did the fine-tuned agent reduce failures on BrowseComp-Plus?

Yes. AWS reported a decline from 22.89% for the base model to 0.68% after fine-tuning.

Sources