When Better Turns Do Not Make Better Agents: Diagnosing the Gap Between Next-Turn Metrics and Workflow Success Md Tahmid Rahman Laskar* , Xue-Yong Fu, Gundeep Singh, Karol Chang, Kevin Sanders, Shi Zong, Tania Habib, Julien Bouvier Tremblay, Shayna Gardiner, Harsh Saini, Matthias Lee, Elena Khasanova, Quinten McNamara, Shashi Bhushan TN (Bold in Author Names Denotes Equal Contribution) Dialpad Inc. Abstract history precedes every decision and compares the model’s prediction with a reference using a lexical Agent models are frequently evaluated one de- arXiv:2609.21187v1 [cs.CL] 18 Sep 2026 cision at a time, where the model predicts the or semantic similarity metric (Lin, 2004; Zhang next action based on the gold interaction his- et al., 2020; Zha et al., 2023), remains a dominant tory, which is scored against a reference. We protocol in practice since it is cheap, reproducible, investigate whether improvement under this and verifiable. protocol is predictive of improved autonomous However, next-turn evaluation has two important workflow execution. We study pre-SFT and su- limitations. First, because each prediction is based pervised fine-tuned (SFT) Qwen3 models at 4B on the correct history, earlier mistakes cannot affect and 14B parameters and Gemma 3 models at 4B and 12B parameters on multi-turn customer- later decisions. It therefore measures how well a support workflows. We find that SFT consis- model responds from a correct state, rather than tently improves text-turn success, and that over- whether it can build and maintain that state during all next-turn success increases for every model an interaction. Second, reference-based metrics under gold-history evaluation. However, these may reward an incorrect response that resembles improvements do not transfer to autonomous the reference while penalizing a valid response ex- workflow execution. Tool-specific gains also pressed differently. This concern has also been vary across metrics and models. None of the noted by Alkhouli et al. (2025) in their CONFETTI four SFT models succeeds under holistic work- flow evaluation, with strict trajectory comple- benchmark, in which they mention that the gold tion reaching at most 10.4% workflow success. trajectory could “artificially inflate” later-turn per- Our results show that next-turn evaluation is formance through in-context learning. Nonetheless, not a reliable proxy for workflow success, mo- they did not compare the turn-level evaluation in tivating separate reporting of text quality, local CONFETTI with end-to-end autonomous execu- action correctness, tool execution, and end-to- tion, where the model must continue from its own end task completion. previous decisions. 1 Introduction To this end, our study directly measures this gap using multi-turn customer-support workflows in Language-model agents deployed in production the task-oriented dialogue evaluation setting (Qin settings must produce more than just plausible re- et al., 2023; Budzianowski et al., 2018). Our re- sponses. For instance, a workflow agent needs search question is whether improvements in gold- to generate a factually correct answer, follow the history next-action scores predict improvements domain policy, invoke tools with state-dependent in end-to-end workflow success. We compare pre- arguments (Schick et al., 2023; Qin et al., 2024), SFT and supervised fine-tuned (SFT) (Wei et al., gather missing information, incorporate tool results, 2022; Ouyang et al., 2022) Qwen3 models (Yang and recover from failed actions. An error early in et al., 2025) at 4B and 14B parameters and Gemma an interaction changes the context for every later 3 models (Gemma Team, 2025) at 4B and 12B. decision, so agent quality is fundamentally a prop- Under gold-history evaluation, SFT improves text erty of a trajectory, not the average quality of its quality and overall turn-level performance across turns, a premise underlying recent interactive agent all model scales. However, when models must benchmarks (Yao et al., 2024; Barres et al., 2025). complete workflows using their own previous out- Nevertheless, next-turn evaluation, where the gold puts, nearly all trajectories requiring tool use led * Corresponding Author: tahmid.rahman@dialpad.com to failure. Thus, the same SFT models that appear effective under turn-level evaluation are found to teractive environments (Liu et al., 2024), including be ineffective in end-to-end workflow execution. policy-constrained tool–agent–user interactions in In this paper, we investigate three research ques- τ -bench and dual-control interaction in τ 2 -bench tions : (i) whether turn-level improvements predict (Yao et al., 2024; Barres et al., 2025). end-to-end workflow success, (ii) whether aggre- We complement these benchmarks by holding gate scores conceal differences between text gener- workflows and models fixed while varying only ation and tool execution, and (iii) how early errors the evaluation protocol, isolating how gold-history, affect later decisions. Across multi-turn workflows, strict tool-matching, and closed-loop evaluation we find that SFT consistently improves gold-history support different conclusions for identical outputs. turn-level performance but rarely enables success- Supervised fine-tuning for agent behavior. Su- ful execution of tool-requiring workflows. Sep- pervised fine-tuning on instruction and dialogue arating response generation from tool use further data is a standard method for adapting models to shows that improvements in one capability can hide follow instructions and use tools (Wei et al., 2022; failures in the other, while trajectory analysis re- Ouyang et al., 2022; Wang et al., 2023; Qu et al., veals error propagation that gold-history evaluation 2025), and compact open-weight models are an cannot capture. These findings motivate a mini- increasingly attractive target for this adaptation in mal reporting framework that combines turn-level cost-sensitive deployments (Fu et al., 2024). We scores with end-to-end success and trajectory-level treat SFT as a fixed, realistic intervention and ask error analysis. whether its gains, measured under a gold-history protocol, transfer to closed-loop execution of the 2 Related Work same workflows. Text-generation and reference-based metrics. 3 Experimental Setup Reference-based metrics such as ROUGE mea- sure lexical overlap (Lin, 2004), while BERTScore 3.1 Data and Models and AlignScore measure embedding similarity and Training data. We collect a proprietary dataset source–candidate factual alignment, respectively from Dialpad1 covering 130 customer-support (Zhang et al., 2020; Zha et al., 2023). These are workflows that are constructed from business con- well suited to isolated response turns but do not di- versations (Fu et al., 2022; Laskar et al., 2023; rectly measure policy compliance or tool execution. Khasanova et al., 2025; Laskar et al., 2025b). Each LLM judges can recognize a valid response that workflow is represented by the following: a domain differs from the reference, but their conclusions policy, a set of available tool schemas, and one or remain sensitive to the rubric and context supplied more user goals. Following prior work on tool-use to the judge (Laskar et al., 2025a; Gu et al., 2026). scenario generation and simulated users (Li et al., Tool-using language agents. A separate line of 2023; Qin et al., 2024; Yao et al., 2024), we con- work studies models that select and invoke exter- struct scenarios that combine these elements and nal tools, typically evaluating tool generalization generate multi-turn conversations with two sepa- across a broad catalog of functions (Schick et al., rate instances of GPT-5 (Singh et al., 2025): one 2023; Yao et al., 2023; Li et al., 2023; Qin et al., acts as the user and the other as the workflow agent. 2024; Laskar et al., 2026), with CONFETTI extend- The resulting trajectories contain user messages, ing this to conversational, turn-level function call- assistant responses, tool calls, and tool results. ing (Alkhouli et al., 2025). Our setting instead fixes At first, we generated 1800 conversations. the tool catalog and governing policy per workflow Then, each generated conversation is independently in advance and asks whether an agent can execute checked by Claude-4.5-Opus2 and Gemini-2.5-Pro the full multi-turn trajectory, including recovering (Comanici et al., 2025) for logical consistency, from errors. missing workflow steps, and policy compliance; Task-oriented dialogue and interactive agent a conversation is retained only when both judges benchmarks. Task-oriented dialogue has long approve it. This filtering removes 773 conversa- studied goal completion through multi-turn inter- tions and leaves 1,027 validated conversations. We action and API-grounded slot filling (Qin et al., 1 https://www.dialpad.com/ 2023; Budzianowski et al., 2018; Rastogi et al., 2 https://www.anthropic.com/news/claude-opus-4 2020). Recent benchmarks extend evaluation to in- -5 Figure 1: An Overview of our Evaluation Protocol then convert every assistant decision into a next- sponses and tool calls) based on the given policy, action example: the input contains the policy, tools, tool definitions, gold history, and reference action. and preceding history, and the target is either a The judge model is required to return a binary suc- natural-language response or a structured tool call. cess label by assessing action correctness while The final training split contains 5,834 examples adhering to the domain policy. (4,093 text responses and 1,741 tool calls), with (iii) Strict tool correctness. For the 166 reference 664 additional validation examples (466 text and tool-call decisions, we write a deterministic pars- 198 tool). These generated conversations contain ing script that requires an exact function-name and 13 turns on average. normalized-argument match between the predic- Evaluation data. The held-out evaluation split con- tion and reference. We report exact accuracy and tains 84 validated conversations, yielding 542 next- argument F1, with no partial credit for a plausible action examples: 376 natural-language responses but non-matching action. and 166 tool calls. Conversations contain 2–18 as- (iv) Closed-loop replay. We replay all 84 conver- sistant decisions (median 6). The split is held out at sations using each model’s own generated assistant the conversation level, and no training conversation history rather than gold history, with a determinis- snippet is reused in evaluation. Gold-history met- tic user simulator supplying user turns. We use rics score the 542 examples independently; closed- a deterministic parser to evaluate the tool calls. loop evaluation instead assesses each model across A correctly matched tool call receives the state- the 84 complete conversations by leveraging their dependent result, while an invalid call receives an state-dependent tool results. error and up to two retries. The workflow succeeds Models. We compare public Qwen3 models (Yang only if the model reaches the end of the conversa- et al., 2025) at 4B and 14B parameters and Gemma tion without any unresolved tool errors. 3 instruction-tuned models (Gemma Team, 2025) (v) Holistic workflow judgment. Gemini-2.5-Pro at 4B and 12B (P RE -SFT) with the corresponding judges every closed-loop replay for full user-goal models after full supervised fine-tuning (SFT). Ev- completion where successful workflows do not ery model uses the same training and validation have any policy violation and no tool errors. examples and the same five-epoch recipe. 4 Results 3.2 Evaluation Protocols Table 1 reports performance across the evaluation We evaluate every model under five protocols, each protocol, from gold-history turn-level evaluation targeting a different capability along the path from to end-to-end workflow execution. Overall, SFT producing a plausible utterance to completing an substantially improves text-based responses, but autonomous workflow (see Figure 1). these gains transfer weakly to tool execution and (i) Text similarity. For every natural-language rarely yield completed workflows. decision, the model receives the gold history pre- ceding it, the policy, and the tool definitions, and 4.1 SFT improves text-based turn-level we compare its response with the reference using performance the ROUGE-1 metric. Under gold-history evaluation, SFT improves per- (ii) Gold-history turn success. Gemini-2.5-Pro formance across all model families and scales. Av- at temperature zero judges all model-generated re- eraged across the four models, ROUGE-1 increases sponses (542 turns, covering both text-based re- by 24.4 points, from 24.8 to 49.2, while LLM- Qwen3 Gemma 3 4B 14B 4B 12B Level Metric Scoring ZS SFT ZS SFT ZS SFT ZS SFT Text similarity ROUGE-1 Deterministic 30.6 49.8 12.9 50.5 25.6 46.9 30.0 49.4 Gold-history turn All-turn success LLM judge 29.2 50.7 37.1 54.6 20.1 41.1 27.5 44.5 Gold-history turn Text-turn success LLM judge 32.4 61.4 41.8 64.4 24.7 50.8 33.8 56.9 Gold-history turn Tool-turn success LLM judge 21.7 26.5 26.5 32.5 9.6 19.3 13.3 16.3 Strict tool check Exact call accuracy Deterministic 3.6 13.9 5.4 18.1 3.0 0.0 1.8 1.8 Strict tool check Argument F1 Deterministic 6.3 24.8 10.7 29.0 5.1 0.5 4.5 2.9 Closed-loop replay Tool-workflow completion Deterministic 0/77 3/77 0/77 8/77 0/77 0/77 0/77 0/77 Holistic judge Tool-workflow success LLM judge 0/77 0/77 0/77 0/77 0/77 0/77 0/77 0/77 Table 1: Results across the evaluation protocols for Qwen3 (4B/14B) and Gemma 3 (4B/12B) under Zero-Shot (ZS) and Supervised Fine-Tuning (SFT) settings. judged text-turn success rises by 25.2 points, from policy violations, or improper use of tool results. 33.2% to 58.4%. Consequently, all-turn success improves by 19.3 points, from 28.5% to 47.7%. 4.4 Evaluation protocol changes conclusion These consistent gains show that SFT helps mod- The two protocols yield different conclusions about els better predict the expected next response when the same SFT interventions. Gold-history evalu- given the correct interaction history. ation indicates broad improvement: every model 4.2 Turn-level gains fail to extend to tool use achieves higher ROUGE-1, text-turn success, and all-turn success after SFT. End-to-end execution Tool-specific improvements are substantially instead shows that nearly all tool-requiring work- smaller. LLM-judged tool-turn success increases flows still fail, with none succeeding under holis- by only 5.9 points on average, from 17.8% to tic evaluation. This discrepancy arises because 23.7%, compared with 25.2 points for text turns. gold-history evaluation restores the correct state Exact call accuracy increases from 3.5% to 8.5%, before each decision. It tests whether a model and argument F1 from 6.7% to 14.3%. These gains can predict the next action from a correct history, are concentrated in Qwen3-4B and Qwen3-14B. but not whether it can construct and maintain that Both Gemma 3 models show little or negative im- history through its own decisions. End-to-end ex- provement; for example, Gemma 3-4B declines ecution exposes this limitation by allowing early from 3.0% to 0.0% in exact call accuracy and from errors to propagate. Turn-level evaluation is there- 5.1% to 0.5% in argument F1. Thus, turn-level fore informative but measures a narrower capability scores based on gold history can obscure poor tool than workflow execution. Reliable agent evaluation use: frequent text turns improve strongly after SFT should report text quality, tool correctness, work- and raise the overall score even when exact tool flow completion, and holistic trajectory success execution remains poor or deteriorates. separately, rather than treating aggregate next-turn 4.3 Gold-history overestimates workflow performance as evidence of end-to-end capability. success 5 Conclusion The gap widens when models execute workflows using their own prior outputs. Before SFT, no Across four model pairs from two families, SFT model completes a tool-requiring workflow. After clearly improves reference-matched text and next- SFT, Qwen3-4B and Qwen3-14B complete only action prediction under gold-history evaluation. 3 and 8 workflows, respectively, while neither However, tool-call correctness varies across met- Gemma 3 model completes any. The best result, rics, model sizes, and families, while end-to-end from Qwen3-14B, is only 10.4% workflow com- completion remains low as holistic evaluation finds pletion. Moreover, deterministic completion does no successful tool-requiring workflow at any scale. not imply overall interaction success. The holistic These results show that strong next-turn perfor- judge identifies no successful tool-requiring work- mance does not necessarily indicate that an agent flow for any model, before or after SFT. A work- can complete a workflow using its own interac- flow may complete its tool calls while still failing tion history. Future work should validate this find- due to missing information, incorrect responses, ing across additional domains, model families, and interactive environments, while developing train- transcripts. In Proceedings of the Eighth Workshop ing methods that explicitly target error recovery, on Noisy User-generated Text (W-NUT 2022), pages 96–100. state maintenance, and end-to-end task comple- tion. More broadly, agent evaluations should com- Xue-Yong Fu, Md Tahmid Rahman Laskar, Elena bine turn-level metrics with tool correctness and Khasanova, Cheng Chen, and Shashi Bhushan Tn. workflow-level success to provide a more complete 2024. Tiny titans: Can smaller large language mod- account of agent capability. els punch above their weight in the real world for meeting summarization? In Proceedings of the 2024 Conference of the North American Chapter of the Limitations Association for Computational Linguistics: Human Language Technologies (Volume 6: Industry Track), Our evaluation covers two model families and one pages 387–394. proprietary customer-support domain, limiting gen- eralizability. Closed-loop replay requires exact ref- Gemma Team. 2025. Gemma 3. arXiv preprint erence tool calls and uses deterministic gold user arXiv:2503.19786. turns, potentially rejecting valid alternatives and Jiawei Gu, Xuhui Jiang, Zhichao Shi, Hexiang Tan, underrepresenting real interactions. Xuehao Zhai, Chengjin Xu, Wei Li, Yinghan Shen, Shengjie Ma, Honghao Liu, and 1 others. 2026. A Ethics Statement survey on llm-as-a-judge. The Innovation, 7(6). All conversations are synthetic and contain no cus- Elena Khasanova, Harsh Saini, Md Tahmid Rah- tomer data or personally identifiable information. man Laskar, Xue-Yong Fu, Cheng Chen, and To help facilitate future work, sanitized prompt tem- Shashi Bhushan Tn. 2025. Dacip-rc: Domain adap- tive continual instruction pre-training via reading plates are provided in the Appendix (see Section A comprehension on business conversations. In Pro- and Section B). ceedings of the 2025 Conference on Empirical Meth- ods in Natural Language Processing: Industry Track, pages 1867–1877. References Md Tahmid Rahman Laskar, Cheng Chen, Xue-yong Fu, Tamer Alkhouli, Katerina Margatina, James Gung, Mahsa Azizi, Shashi Bhushan, and Simon Corston- Raphael Shu, Claudia Zaghi, Monica Sunkara, and Oliver. 2023. Ai coach assist: An automated ap- Yi Zhang. 2025. Confetti: Conversational function- proach for call recommendation in contact centers calling evaluation through turn-level interactions. In for agent coaching. In Proceedings of the 61st An- Proceedings of the 63rd Annual Meeting of the As- nual Meeting of the Association for Computational sociation for Computational Linguistics (Volume 1: Linguistics (Volume 5: Industry Track), pages 599– Long Papers), pages 7993–8006. 607. Victor Barres, Honghua Dong, Soham Ray, Xujie Si, Md Tahmid Rahman Laskar, Xue-Yong Fu, and Karthik Narasimhan. 2025. τ 2 -bench: Evaluat- Seyyed Saeed Sarfjoo, Quinten McNamara, ing conversational agents in a dual-control environ- Jonas Robertson, and Shashi Bhushan TN. 2026. ment. arXiv preprint arXiv:2506.07982. From text to voice: A reproducible and verifiable Paweł Budzianowski, Tsung-Hsien Wen, Bo-Hsiang framework for evaluating tool calling llm agents. Tseng, Iñigo Casanueva, Stefan Ultes, Osman Ra- arXiv preprint arXiv:2605.15104. madan, and Milica Gašić. 2018. MultiWOZ: A large- scale multi-domain wizard-of-oz dataset for task- Md Tahmid Rahman Laskar, Mohammed Saidul Islam, oriented dialogue modelling. In Proceedings of the Ridwan Mahbub, Ahmed Masry, Mizanur Rahman, 2018 Conference on Empirical Methods in Natural Amran Bhuiyan, Mir Tafseer Nayeem, Shafiq Joty, Language Processing (EMNLP), pages 5016–5026. Enamul Hoque, and Jimmy Xiangji Huang. 2025a. Judging the judges: Can large vision-language mod- Gheorghe Comanici, Eric Bieber, Mike Schaekermann, els fairly evaluate chart comprehension and reason- Ice Pasupat, Noveen Sachdeva, Inderjit Dhillon, Mar- ing? In Proceedings of the 63rd Annual Meeting of cel Blistein, Ori Ram, Dan Zhang, Evan Rosen, and the Association for Computational Linguistics (Vol- 1 others. 2025. Gemini 2.5: Pushing the frontier with ume 6: Industry Track), pages 1203–1216. advanced reasoning, multimodality, long context, and next generation agentic capabilities. arXiv preprint Md Tahmid Rahman Laskar, Julien Bouvier Tremblay, arXiv:2507.06261. Xue-Yong Fu, Cheng Chen, and Shashi Bhushan Tn. 2025b. Ai knowledge assist: An automated approach Xue-Yong Fu, Cheng Chen, Md Tahmid Rahman Laskar, for the creation of knowledge bases for conversa- Shashi Bhushan Tn, and Simon Corston-Oliver. 2022. tional ai agents. In Proceedings of the 2025 Con- An effective, performant named entity recognition ference on Empirical Methods in Natural Language system for noisy business telephone conversation Processing: Industry Track, pages 1856–1866. Minghao Li, Yingxiu Zhao, Bowen Yu, Feifan Song, Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Hangyu Li, Haiyang Yu, Zhoujun Li, Fei Huang, Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh and Yongbin Li. 2023. Api-bank: A comprehensive Hajishirzi. 2023. Self-instruct: Aligning language benchmark for tool-augmented llms. In Proceedings models with self-generated instructions. In Proceed- of the 2023 conference on empirical methods in natu- ings of the 61st Annual Meeting of the Association ral language processing, pages 3102–3116. for Computational Linguistics (ACL), pages 13484– 13508. Chin-Yew Lin. 2004. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Branches Out, pages 74–81. Guu, Adams Wei Yu, Brian Lester, Nan Du, An- drew M. Dai, and Quoc V. Le. 2022. Finetuned lan- Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu guage models are zero-shot learners. In International Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Conference on Learning Representations (ICLR). Men, Kejuan Yang, Shudan Zhang, Xiang Deng, Ao- han Zeng, Zhengxiao Du, Chenhui Zhang, Sheng An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Shen, Tianjun Zhang, Yu Su, Huan Sun, and 3 others. Binyuan Hui, Bo Zheng, Bowen Yu, Chang 2024. AgentBench: Evaluating LLMs as agents. In Gao, Chengen Huang, Chenxu Lv, and 1 others. International Conference on Learning Representa- 2025. Qwen3 technical report. arXiv preprint tions (ICLR). arXiv:2505.09388. Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Shunyu Yao, Noah Shinn, Pedram Razavi, and Karthik Carroll Wainwright, Pamela Mishkin, Chong Zhang, Narasimhan. 2024. τ -bench: A benchmark for tool- Sandhini Agarwal, Katarina Slama, Alex Ray, and 1 agent-user interaction in real-world domains. arXiv others. 2022. Training language models to follow in- preprint arXiv:2406.12045. structions with human feedback. Advances in neural information processing systems, 35:27730–27744. Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2023. Libo Qin, Wenbo Pan, Qiguang Chen, Lizi Liao, Zhou ReAct: Synergizing reasoning and acting in language Yu, Yue Zhang, Wanxiang Che, and Min Li. 2023. models. In International Conference on Learning End-to-end task-oriented dialogue: A survey of tasks, Representations (ICLR). methods, and future directions. In Proceedings of the 2023 Conference on Empirical Methods in Natural Yuheng Zha, Yichi Yang, Ruichen Li, and Zhiting Hu. Language Processing, pages 5925–5941. 2023. AlignScore: Evaluating factual consistency with a unified alignment function. In Proceedings Yujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu, Lan of the 61st Annual Meeting of the Association for Yan, Yaxi Lu, Yankai Lin, Xin Cong, Xiangru Tang, Computational Linguistics (Volume 1: Long Papers), Bill Qian, and 1 others. 2024. Toolllm: Facilitating pages 11328–11348. large language models to master 16000+ real-world apis. In International Conference on Learning Rep- Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. resentations, volume 2024, pages 9695–9717. Weinberger, and Yoav Artzi. 2020. BERTScore: Evaluating text generation with BERT. In Inter- Changle Qu, Sunhao Dai, Xiaochi Wei, Hengyi Cai, national Conference on Learning Representations Shuaiqiang Wang, Dawei Yin, Jun Xu, and Ji-Rong (ICLR). Wen. 2025. Tool learning with large language mod- els: A survey. Frontiers of Computer Science, 19(8):198343. A Sanitized Data-Generation Prompts Abhinav Rastogi, Xiaoxue Zang, Srinivas Sunkara, The following prompts preserve the roles, inputs, Raghav Gupta, and Pranav Khaitan. 2020. Towards scalable multi-domain conversational agents: The and decisions used in the pipeline while replacing schema-guided dialogue dataset. In Proceedings of proprietary policies, tools, and company informa- the AAAI Conference on Artificial Intelligence, pages tion with placeholders. 8689–8696. A.1 Scenario Generation Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Given the domain policy and available tools, create a Raileanu, Maria Lomeli, Luke Zettlemoyer, Nicola realistic customer-support scenario for this workflow. Cancedda, and Thomas Scialom. 2023. Toolformer: Language models can teach themselves to use tools. Specify: - the user's objective and relevant initial state; In Advances in Neural Information Processing Sys- - the desired final state; tems (NeurIPS). - information the user initially knows; - the expected tool operations; and - observable criteria for successful completion. Aaditya Singh, Adam Fry, Adam Perelman, Adam Tart, Adi Ganesh, Ahmed El-Kishky, Aidan McLaughlin, The scenario must be solvable using only the policy and Aiden Low, AJ Ostrow, Akhila Ananthram, and 1 oth- tools supplied below. Use synthetic identities and data. ers. 2025. Openai gpt-5 system card. arXiv preprint DOMAIN POLICY: arXiv:2601.03267. B Sanitized Evaluation-Judge Prompts AVAILABLE TOOLS: These prompts show the information supplied to the judges and the binary decisions reported in A.2 GPT-5 User Simulator the paper. Auxiliary diagnostic fields used during Act as the customer described below. Remain in character analysis are omitted. and pursue the assigned objective naturally. - Reveal information gradually and only when appropriate. B.1 Gold-History Turn Judge - Answer the agent's questions consistently with the scenario state and persona. Evaluate one assistant decision in a multi-turn - Do not mention tools, policies, prompts, or simulation. customer-support workflow. Judge only what should happen - End when the goal is met or cannot be completed. at this turn; do not require final workflow completion at an intermediate turn. SCENARIO AND OBJECTIVE: Use the policy, tools, correct history, and reference action to determine whether the prediction is an SYNTHETIC IDENTITY AND PERSONA: acceptable next action. Alternative wording and safe, semantically equivalent actions are allowed. A tool call must select an appropriate function and use arguments CONVERSATION SO FAR: grounded in the supplied history. DOMAIN POLICY: Return only the next user message. AVAILABLE TOOLS: A.3 GPT-5 Workflow Agent CORRECT HISTORY BEFORE THIS TURN: Act as a customer-support agent operating under the domain policy below. REFERENCE ACTION: - Follow the workflow and all policy constraints. - Maintain the conversation state and request missing information when needed. MODEL PREDICTION: - Call an available tool when the workflow requires it. - Use tool results before making state-dependent claims. - Handle invalid requests and tool errors safely. Return JSON only: { Return either the next natural-language response or a "turn_success": true or false, structured tool call. "failure_category": "none or primary failure", "rationale": "brief evidence-based explanation" DOMAIN POLICY: } AVAILABLE TOOLS: B.2 Holistic Workflow Judge CONVERSATION SO FAR: Evaluate whether the candidate replay completes the user's workflow under the supplied domain policy. Use the reference trajectory to understand the scenario, required state changes, and acceptable final outcome. Do A.4 Conversation Validation not reward wording similarity. Safe alternative wording or an equivalent valid path is allowed, but required This prompt is run independently with each valida- verification, policy constraints, and consequential tool actions must be satisfied. tion judge. A conversation is retained only when both judges approve it. Accepted tool attempts received the recorded tool result. Rejected attempts were invalid at that workflow state. A replay that ends before the user goal is resolved is not Review the complete synthetic conversation for use as a successful. workflow-training example. DOMAIN POLICY AND TOOLS: Approve it only if the agent follows the supplied policy, uses tools appropriately, remains consistent with tool results, and reaches a sensible outcome for the scenario. REFERENCE TRAJECTORY: Reject it for a missing required step, contradiction, unsupported claim, invalid tool use, policy violation, or inconsistent simulated-user behavior. CANDIDATE CLOSED-LOOP REPLAY: DOMAIN POLICY AND TOOLS: Return JSON only: { SCENARIO: "success": true or false, "failure_category": "none or primary failure", "rationale": "brief evidence-based explanation" GENERATED CONVERSATION: } Return JSON only: { "decision": "APPROVE or REJECT", "reasons": ["brief reason"] }