Salesforce Koa: An Enterprise Language Model for Agentic Tool Use Zixiang Chen* , Sufeng Niu* , Yingchi Liu* , Wenting Zhao* , Akshara Prabhakar* , Shubham Mehrotra* Bin Bi, Zhujun Lan, Katherine Tan, Mohammad Ramezanali, Tulika Manoj Awalgaonkar, Monojit Banerjee, Jielin Qiu, Shiva Kumar Pentyala, Zhepeng Cen, Anupam Tripathi, Ali Ziaei, Regunathan Radhakrishnan Darvish Lee Shadravan, Shelby Heinecke, Sitaram Asur, Silvio Savarese, James Zhu, Phil Mui† , Huan Wang† arXiv:2609.15066v1 [cs.CL] 14 Sep 2026 Salesforce Agentforce & AI Research September 15, 2026 Abstract We present Salesforce Koa, an enterprise language model built by post-training the open-weight Nemotron-3-Super-120B foundation model with reinforcement learning using Group Relative Policy Optimization (GRPO). Salesforce Koa is trained on public and synthetically generated data, with no customer data, to improve tool use and agentic capabilities while preserving strong general-purpose performance. Its distinctive component is a simulation-to-reward pipeline that expands workflow specifi- cations into persona-conditioned multi-turn tasks with task-resolution rewards grounded in successful tool use for data-dependent requests. For enterprise domains, these specifications are written in Agent Script, Salesforce’s declarative language for building Agentforce agents; for public tool-use domains, we synthesize the workflow structure directly. The same simulation and grounded-reward machinery drives GRPO across both. Across public tool-use, agentic-reasoning, and enterprise Customer Relationship Man- agement (CRM) benchmarks, Salesforce Koa improves over its open-weight base, with the clearest gains on multi-turn tool use, and surpasses a strong proprietary baseline while remaining below the strongest frontier models. These results show that specification-driven reinforcement learning is a practical path to specializing open-weight foundation models for enterprise agentic tasks. 1 Introduction Large language models (LLMs) increasingly power enterprise AI systems, where, unlike general-purpose assistants, they must invoke external tools, retrieve structured information, and complete multi-step workflows across business applications. Robust tool use and agentic reasoning are therefore essential enterprise capabilities. Many such systems rely on proprietary frontier models, yet enterprises often need more control over deployment, customization, governance, and cost than a single external provider allows. Recent open- weight foundation models such as Llama [3], Gemma [4], Qwen [18], Mistral [7], and Nemotron [10] have closed much of the public-benchmark gap, and access to weights lets organizations post-train for their own domains. A central question is whether they can be effectively specialized for enterprise agentic tasks while preserving general-purpose capability. LLMs are increasingly deployed as tool-using agents [17, 12, 15], with benchmarks such as BFCL [13] and CRMArena [6] establishing tool use and agentic reasoning as critical capabilities. Post-training is the standard route: supervised fine-tuning adapts foundation models to downstream tasks [11], while reinforcement learning, via RLHF, Constitutional AI, and Group Relative Policy Optimization (GRPO) [5], further improves reasoning and decision-making. Most of this work targets general-purpose reasoning over * † Equal contribution (co-first authors). Co-corresponding authors. 1 S ALESFORCE KOA T ECHNICAL R EPORT public tools; in contrast, enterprise applications require operating over organization-specific schemas, APIs, and policies, the setting we study here. We introduce Salesforce Koa, an enterprise language model built by post-training the open-weight Nemotron-3-Super-120B foundation model (Nemotron-120B) with GRPO, using only public and synthetically generated data (no customer data). The distinctive component of our pipeline is specification-driven task construction: for enterprise domains, we build training tasks from declarative agent specifications written in Agent Script [16], Salesforce’s declarative language for Agentforce agents. A specification captures an agent’s routing structure, subagents, typed actions, tool scopes, and workflow instructions; a simulation pipeline expands it into scenario- and persona-conditioned tasks that NeMo Gym executes as online rollouts, with grounded task-resolution rewards driving GRPO. This creates a direct bridge between agent authoring and model post-training: the same specifications that configure an agent also structure its rollout tasks and resolution criteria. Our experiments show that across public tool-use, agentic-reasoning, and enterprise CRM benchmarks Salesforce Koa improves over its open-weight base, with the clearest gains on multi-turn tool use, and surpasses a strong proprietary baseline (GPT-4.1) while remaining below the strongest frontier models. Together these results demonstrate that spec-driven RL is a practical path to specializing open-weight foundation models for enterprise agentic tasks. We additionally report a scoped SFT-vs-RL comparison (Appendix J): from our already RL-post-trained base, RL improves multi-turn tool use substantially while SFT adds little. We treat this as an observation specific to our starting point rather than a general claim. 2 Methodology Salesforce Koa is an enterprise language model built by post-training the open-weight Nemotron-120B foundation model with reinforcement learning using Group Relative Policy Optimization (GRPO). Training uses only public resources and synthetically generated interactions; no customer data is used. The distinctive component of our pipeline is specification-driven task construction: declarative enterprise agent specifications are expanded into executable, persona-conditioned multi-turn environments whose rewards are grounded in successful tool use. We describe environment and task construction (Section 2.1), online rollout and the grounded reward (Section 2.2), and policy optimization (Section 2.3). We apply RL directly to the base model rather than to an SFT checkpoint; a preliminary SFT study that motivated this choice, together with distributed-training and evaluation details, is deferred to Appendix C and G. 2.1 Spec-driven environment and task construction For enterprise domains, we author workflow structure in Agent Script. A specification defines a router, specialized subagents, typed actions, and natural-language reasoning instructions, which together determine how requests are routed, which tools are exposed within each topic, and how each topic handles its workflow. A static extraction pass compiles this specification into a typed workflow graph capturing per-tool argument schemas, declared state effects, routing conditions, and loop/termination configuration; this graph instantiates the executable environment, including the simulated state that tools read and mutate. For public tool-use domains, where no Agent Script specification exists, we instead synthesize the workflow structure directly, producing the same typed workflow graph of tools, argument schemas, and termination conditions; both sources therefore feed the shared simulation and reward machinery below equally. A simulation pipeline expands each specification into scenario- and persona-conditioned sessions: the scenario determines the workflows and user needs exercised, while the persona conditions the simulated customer’s behavior. These sessions are converted into RL examples that begin either from the opening request or from a generated dialogue prefix, exposing the policy to decisions at different depths of a multi-turn interaction. Each example uses a common schema in which agent_ref selects the environment and the remaining fields provide the initial rollout context (task intent, persona, dialogue state, available tools); the full trajectory is generated online. When multiple environments are trained jointly, per-row agent_ref dispatch lets a single run load them all, and differential exposure is realized by repeating an environment’s examples 2 S ALESFORCE KOA T ECHNICAL R EPORT before shuffling. NeMo Gym then executes each example as an online rollout, so workflow authoring and task construction are cleanly separated from rollout execution and policy optimization. 2.2 Online rollout and grounded reward Each example is executed inside a NeMo Gym environment. For simulated-dialogue environments a single frozen helper model plays three roles in every rollout: a customer simulator (produces the next user turn from the sampled persona), a tool/function emulator (produces tool outputs when no production backend exists), and a coverage judge (scores resolution to form the reward). The helper is a separate inference service and is never updated by policy optimization. Appendix B (Figure 1) traces one rollout end to end, making explicit which components are frozen versus trained, the order of interactions within a turn, and where the scalar reward comes from. The reward is built from the coverage judge. The task intent is represented as numbered sub-questions, and the judge emits one resolved/not-resolved verdict per sub-question over the full transcript. The coverage rate is #{sub-questions judged resolved} cov = ∈ [0, 1], (1) #{sub-questions} where missing or malformed verdicts default to not resolved. The scalar reward applies a duplicate-call gate, ( 0, a consecutive duplicate tool call occurred, R= (2) cov, otherwise, and the judge prompt requires that any sub-question needing customer- or system-specific data be marked resolved only when the answer is grounded in a successful relevant tool call; generic factoids may resolve without a tool. The deterministic sandbox environment instead sets R = 1 exactly when the predicted actions reproduce the reference end state and R = 0 otherwise. A few properties of these environments are central to their behavior. The tool list is per rollout, not global: each conversation advertises only the tools available in its source session, and in routed environments the visible set is further restricted to the active subagent, so the policy sees only in-scope tools at any moment. Task intent is constructed deterministically at build time (a numbered list of sub-questions), directly connecting task construction to the reward. A persona is sampled per conversation and held constant, controlling how the simulated user reacts, escalates, or declares resolution. Focus system prompts are not restricted to the conversation start: in routed environments each subagent switch appends a fresh focus/procedure message, re-focusing the policy exactly when the topic changes. Environments differ mainly in interaction topology and verification: a flat customer-support environment exposes all advertised tools to a single agent; a routed customer-service environment adds a router and specialized subagents reached through go_to_* with a focus message on each switch; and a deterministic tool sandbox runs stateful tools with no simulated customer or helper, scored by binary state-equivalence rather than the coverage judge (Appendix A, Table 2). 2.3 Optimization We maximize the expected trajectory reward J (θ) = Ex∼D Eτ ∼πθ (·|x,env) [R(τ )] with GRPO. For each prompt the policy samples a small group of complete trajectories with colocated inference; dynamic sampling discards zero-variance groups and refills until a full batch is assembled. Advantages are estimated by a leave-one-out group baseline (no value network), and a token-level truncated importance-sampling weight corrects the train/generation log-probability mismatch under a single on-policy update per batch. Trajectories with malformed tool-call or thinking syntax have the offending token advantages set to a fixed negative value, and over-length or log-probability-inconsistent trajectories are masked. The full objective, advantage normalization, and validity constraints are given in Appendix E. A complementary single-step training mode for constrained decision turns, which isolates the rare but decisive turns where the agent must communicate with the user rather than call a tool, is described in Appendix F. 3 S ALESFORCE KOA T ECHNICAL R EPORT Table 1: Performance comparison on Tau2Bench, BFCL, and CRM Bench. The Tau2Bench average is weighted by the number of tasks in each domain (Airline: 50, Retail: 114, Telecom: 114). Tau2Bench BFCL CRM Bench Weighted Weighted Model Airline Retail Telecom Avg. Acc. Topic Func. Text Avg. Claude Opus 4.8 69.0 86.2 64.0 74.00 78.18 0.99 0.83 0.79 0.87 OpenAI GPT-5.5 62.5 81.6 95.8 83.99 67.63 0.99 0.82 0.89 0.90 OpenAI GPT-4.1 56.0 74.1 34.2 54.48 53.96 0.98 0.85 0.60 0.81 Nemotron-3-Super-120B 61.5 79.9 60.5 68.64 64.73 0.97 0.71 0.85 0.84 Salesforce Koa (Ours) 62.0 81.6 60.5 69.41 66.63 0.97 0.77 0.85 0.86 3 Experiments We post-train the Nemotron-3-Super v3 (∼120B) policy on multi-domain agentic tool-calling tasks and evaluate it against strong open and proprietary baselines. The final gated-coverage reward and GRPO recipe were developed on a cheaper Nemotron-3-Nano v3 (∼30B) proxy before being ported to the 120B policy; the reward-engineering ablation and the observation that reward stability is scale-dependent are reported in Appendix I. Setup. Both policies are Nemotron reasoning models run in thinking mode during training. Roll- outs are generated against NeMo Gym resource servers spanning salesforce support, healthcare administration, real estate, and a helper-free workplace assistant sandbox. Each simulated- dialogue environment is driven by a shared Nano-30B helper serving all three environment roles (customer simulator, tool emulator, coverage judge). Training runs on a Slurm cluster of 5×NVIDIA B200 nodes (colocated rollout/training + one helper node); see Appendix G for training and evaluation details. Evaluation and baselines. We compare Salesforce Koa with three proprietary frontier models (Opus-4.8, GPT-5.5, and GPT-4.1) and with its open-weight base, Nemotron-3-Super-120B. We evaluate on two public tool-calling benchmarks and one released enterprise benchmark: Tau2Bench [2] (end-to-end multi-turn customer service across airline, retail, and telecom; GPT-4.1 user simulator, four trials, pass^1 averaged), BFCL [13] (agentic tool use across multi-step calling, web search, memory, and stateful tools), and CRM Bench, covering single-turn Salesforce and Agentforce workflows scored along topic, function-call, and free-text accuracy. All checkpoints are served via vLLM in BF16. Results. As shown in Table 1, Salesforce Koa is competitive across all three benchmarks. On Tau2Bench it reaches a task-weighted average of 69.41, edging its Nemotron base (68.64) and outperforming GPT-4.1 by 14.9 points. On BFCL it achieves 66.63%, improving over the base (64.73%) and well above GPT-4.1 (53.96%), though below the strongest proprietary models. On CRM Bench it scores 0.86 overall, close to Opus-4.8 (0.87) and exceeding both GPT-4.1 (0.81) and its base (0.84); its function-call accuracy of 0.77 improves over the base (0.71). Together these show that spec-driven enterprise RL improves over the open-weight base across public and enterprise benchmarks, most clearly on multi-turn tool use, and surpasses a strong proprietary baseline (GPT-4.1) while remaining below the strongest frontier models. We build Salesforce Koa on the RL-trained checkpoint. Appendix J gives a scoped comparison of SFT-only and RL-only adaptation from the same base, which finds RL substantially stronger on multi-turn tool use while SFT is competitive on single-turn CRM tasks, with the important caveat that our base is itself already RL-post-trained. 4 S ALESFORCE KOA T ECHNICAL R EPORT 4 Conclusion We presented Salesforce Koa, an enterprise language model built by post-training the open-weight Nemotron- 120B foundation model with GRPO reinforcement learning. Across public tool-use, agentic-reasoning, and enterprise CRM benchmarks Salesforce Koa improves over its open-weight base, most clearly on multi-turn tool use, and surpasses a strong proprietary baseline (GPT-4.1) while remaining below the strongest frontier models. Its central contribution is a specification-driven RL pipeline in which the same declarative Agent Script specifications that configure an agent also supply the workflow structure for constructing scenario- and persona-conditioned tasks and grounded task-resolution rewards, directly linking agent authoring to model post-training. A key limitation is that our base is itself already RL-post-trained, so our scoped finding, that RL improves multi-turn tool use substantially while SFT adds little, may not hold from a pre-RL checkpoint. Future work includes co-designing the SFT and RL stages, evaluating from a pre-RL foundation, and extending specification-driven environments to broader enterprise domains and more complex agentic workflows. 5 Acknowledgements This model was trained in partnership with NVIDIA. References [1] NeMo AutoModel: DTensor-native SPMD library for scalable and efficient training. https://github. com/NVIDIA-NeMo/Automodel, 2025–2026. GitHub repository. [2] Victor Barres, Honghua Dong, Soham Ray, Xujie Si, and Karthik Narasimhan. τ 2 -Bench: Evaluating conversational agents in a dual-control environment, 2025. URL https://arxiv.org/abs/2506. 07982. [3] Abhimanyu Dubey et al. The Llama 3 herd of models. arXiv preprint arXiv:2407.21783, 2024. [4] Gemma Team et al. Gemma 3 technical report. arXiv preprint arXiv:2503.19786, 2025. [5] Daya Guo et al. DeepSeek-R1: Incentivizing reasoning capability in LLMs via reinforcement learning. arXiv preprint arXiv:2501.12948, 2025. [6] Kung-Hsiang Huang, Akshara Prabhakar, Sidharth Dhawan, Yixin Mao, Huan Wang, Silvio Savarese, Caiming Xiong, Philippe Laban, and Chien-Sheng Wu. CRMArena: Understanding the capacity of LLM agents to perform professional CRM tasks in realistic environments. In Luis Chiruzzo, Alan Ritter, and Lu Wang, editors, Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 3830–3850, Albuquerque, New Mexico, April 2025. Association for Computational Linguistics. ISBN 979-8-89176-189-6. doi: 10.18653/v1/2025.naacl-long.194. URL https://aclanthology. org/2025.naacl-long.194/. [7] Albert Q. Jiang et al. Mistral 7B. arXiv preprint arXiv:2310.06825, 2023. [8] Zuxin Liu, Thai Hoang, Jianguo Zhang, Ming Zhu, Tian Lan, Shirley Kokane, Juntao Tan, Weiran Yao, Zhiwei Liu, Yihao Feng, Rithesh Murthy, Liangwei Yang, Silvio Savarese, Juan Carlos Niebles, Huan Wang, Shelby Heinecke, and Caiming Xiong. APIGen: Automated pipeline for generating verifiable and diverse function-calling datasets. arXiv preprint arXiv:2406.18518, 2024. URL https: //arxiv.org/abs/2406.18518. 5 S ALESFORCE KOA T ECHNICAL R EPORT [9] Xueyan Niu, Bo Bai, Wei Han, and Weixi Zhang. On the non-decoupling of supervised fine-tuning and reinforcement learning in post-training, 2026. URL https://arxiv.org/abs/2601.07389. [10] NVIDIA et al. Nemotron-4 340B technical report. arXiv preprint arXiv:2406.11704, 2024. [11] Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Leike, and Ryan Lowe. Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems, volume 35, pages 27730–27744, 2022. [12] Shishir G. Patil, Tianjun Zhang, Xin Wang, and Joseph E. Gonzalez. Gorilla: Large language model connected with massive APIs. In Advances in Neural Information Processing Systems, 2024. [13] Shishir G. Patil, Huanzhi Mao, Fanjia Yan, Charlie Cheng-Jie Ji, Vishnu Suresh, Ion Stoica, and Joseph E. Gonzalez. The Berkeley Function Calling Leaderboard (BFCL): From tool use to agentic evaluation of large language models. In Proceedings of the 42nd International Conference on Machine Learning, 2025. URL https://proceedings.mlr.press/v267/patil25a.html. [14] Akshara Prabhakar, Zuxin Liu, Ming Zhu, Jianguo Zhang, Tulika Awalgaonkar, Shiyu Wang, Zhiwei Liu, Haolin Chen, Thai Hoang, Juan Carlos Niebles, Shelby Heinecke, Weiran Yao, Huan Wang, Silvio Savarese, and Caiming Xiong. APIGen-MT: Agentic pipeline for multi-turn data generation via simulated agent-human interplay. arXiv preprint arXiv:2504.03601, 2025. URL https://arxiv. org/abs/2504.03601. [15] Yujia Qin, Shihao Liang, Yining Ye, et al. ToolLLM: Facilitating large language models to master 16,000+ real-world APIs. In International Conference on Learning Representations, 2024. [16] Salesforce. Agent Script: Agentforce developer guide. https://developer.salesforce.com/ docs/ai/agentforce/guide/agent-script.html, 2026. [17] Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. Toolformer: Language models can teach them- selves to use tools. In Advances in Neural Information Processing Systems, 2023. [18] An Yang et al. Qwen2 technical report. arXiv preprint arXiv:2407.10671, 2024. [19] Junkeun Yi, Damon Mosk-Aoyama, Baihe Huang, Ritu Gala, Charles Wang, Sugam Dipak Devare, Khushi Bhardwaj, Abhibha Gupta, Oleksii Kuchaiev, Jiantao Jiao, et al. PivotRL: High accuracy agentic post-training at low compute cost. arXiv preprint arXiv:2603.21383, 2026. 6 S ALESFORCE KOA T ECHNICAL R EPORT A Environment interaction and verification Table 2 summarizes how the training environments (Section 2.2) differ in interaction topology and verification. Table 2: Environment interaction and verification mechanisms. The training environments are listed in the Experiments section. Environment Interaction topology Verification type Flat customer Single agent; all advertised tools directly Per-sub-question LLM coverage with a hard support callable; persona-conditioned simulated duplicate-call gate; the judge prompt enforces customer. successful-tool grounding for data-bound questions. Routed customer Router plus specialized subagents reached Same coverage-based reward and service through go_to_*; out-of-scope calls return a duplicate-call gate, with domain-specific scope wrong-subagent error; a focus system message and grounding rules in the judge prompt. is inserted on each switch. Deterministic tool Single agent over stateful tools; no simulated Binary state-equivalence: reward is one only if sandbox customer and no helper. the state produced by the predicted actions matches the state produced by the reference actions. B Training rollout Figure 1 traces one training rollout (Section 2.2) end to end, making explicit which components are frozen versus trained, the order of interactions within a turn, and where the scalar reward comes from. Frozen helper Policy Env agent Resource (sim / emulator (trainable) + router server / judge) response / tool call / go_to_* dispatch (scope + schema check) emulate tool result synthesized output tool result / observation request next customer turn simulated user message repeat until the session ends or the turn budget is reached coverage judge over sub-questions per-sub-question verdicts scalar reward + diagnostics Figure 1: One training rollout. Solid arrows are the trained/rollout path; dashed arrows are frozen-helper responses. The policy (left) is the only trained component; the reward is produced by the coverage judge at the end. 7 S ALESFORCE KOA T ECHNICAL R EPORT C Preliminary supervised fine-tuning study Before committing to RL, we ran a preliminary SFT study to gauge how much imitation on curated tool-use trajectories could add on top of an already heavily post-trained foundation model. The finding, limited headroom on multi-turn tool use, motivated applying RL directly to the base model (Section 2). Objective and data. SFT is standard behavior cloning: the policy πθ is trained with a masked next-token loss whose mask restricts supervision to target tokens (assistant messages and tool calls) and excludes prompt, system, user, and tool-result tokens. The entire corpus is synthetic, generated with an automated pipeline built on APIGen [8] and the multi-turn APIGen-MT [14], and contains no customer, proprietary, or privatized data. We draw only on publicly available function-calling resources: tool collections from BFCL [13] and Tau2Bench (Airline) [2], used purely as tool schemas and executable environments rather than as conversations. Following APIGen-MT, we instantiate a simulated environment per tool collection and generate complete multi-turn interactions; every trajectory is automatically verified (tool calls must execute successfully and the session must satisfy the task’s success criteria), and we retain only successful, verified trajectories. Per-step supervision. Each session is normalized to a common conversational schema (system, user, assistant, tool-call, tool-result items) with the callable tool schemas attached. For every assistant decision point we form one training example whose input is the conversation prefix and whose target is the reference assistant action, so a single trajectory supervises the policy at every step it must act, including deep in a multi-turn session. Over-long prefixes are dropped with a conservative token estimate, retained examples are rendered with the model’s chat template (so SFT and inference share a surface form), and sessions are split into train/validation before per-turn examples are generated so no validation session leaks through a prefix. Appendix D gives a concrete example. Training procedure. We perform full-parameter fine-tuning of Nemotron-120B using NeMo Auto- model [1], distributed over 32 H200 GPUs (4 nodes × 8) with FSDP2, expert parallelism (ep_size = 8), and activation checkpointing, at a maximum sequence length of 8192 and global batch size 32 (Adam, β = (0.9, 0.999), ϵ = 10−8 , no weight decay). Because the base model has already undergone elaborate post-training, our central concern is avoiding catastrophic forgetting; we therefore use a deliberately conser- vative recipe (peak learning rate 5 × 10−8 , linear warmup, cosine decay toward 10−9 , gradient clipping at 1.0, early stopping). In practice the validation loss moved marginally, an early sign of limited headroom, and we select the lowest-validation-loss checkpoint. Observations. Gains concentrate in single-turn and static agentic categories, not multi-turn tool use (Table 3): on BFCL, live parallel and parallel-multiple AST rise sharply (75.0% → 87.5% and 79.2% → 87.5%), memory improves (52.9% → 57.4%), web search (base) improves (77% → 83%), and relevance detection jumps (68.8% → 87.5%), partly offset by regressions on some non-live AST subsets and irrelevance detection. The BFCL multi-turn category, the capability most central to enterprise agents, is essentially unchanged (54.1% → 53.2%). This suggests that, from an already heavily post-trained base, SFT on our corpus offers limited additional headroom for the target capability, motivating RL applied directly to the base model. We view a combined SFT-then-RL pipeline as promising future work; realizing its benefit is nontrivial, since SFT and RL are not cleanly decoupled in post-training [9]. D SFT training example A single SFT example (fabricated for illustration) has the following shape: a prefix of system/user/tool items as input, and one masked assistant target. 8 S ALESFORCE KOA T ECHNICAL R EPORT { "agent_ref": {"type": "responses_api_agents", "name": "support_agent"}, "responses_create_params": { "input": [ {"type":"message","role":"system", "content":""}, {"type":"message","role":"system", "content":""}, {"type":"message","role":"user","content":"Where is my order A123?"}, {"type":"function_call","call_id":"c1","name":"get_order_status", "arguments":"{\"order_id\":\"A123\"}"}, {"type":"function_call_output","call_id":"c1", "output":"{\"status\":\"shipped\",\"eta_days\":2}"} ], "tools": [ {"type":"function","function":{"name":"get_order_status", "parameters":{"type":"object", "properties":{"order_id":{"type":"string"}}, "required":["order_id"]}}} ] }, // masked SFT target = the next assistant turn: "target": {"type":"message","role":"assistant", "content":"Your order A123 has shipped and should arrive in about 2 days. Anything else?"} } Only the target tokens contribute to the SFT loss. E GRPO optimization details We maximize the expected trajectory reward   J (θ) = Ex∼D Eτ ∼πθ (· | x, env) R(τ ) , (3) where a trajectory τ is a full rollout inside an executable environment and R(τ ) is the reward of Section 2.2. Because the reward is available only at trajectory end and we train no value network, we estimate advantages by comparing several trajectories sampled for the same prompt. Differential exposure across environments (Section 2.1) is realized by repeating an environment’s examples before shuffling, ! (e) ] Dtrain = Shuffle re Dtrain , (4) e∈E with per-environment factors re ; where prefix expansion already multiplied an environment’s examples, its explicit factor is one. The validation corpus is not upsampled, so aggregate validation reward reflects natural environment prevalence. Group sampling and dynamic sampling. For each prompt the policy samples a small group of complete trajectories with colocated inference. Because a group in which every trajectory earns the same reward yields no learning signal, dynamic sampling retains only trajectories whose group reward has non-zero variation and refills across generation batches until a full training batch is assembled. 9 S ALESFORCE KOA T ECHNICAL R EPORT Leave-one-out normalized advantages. For trajectory i in a prompt’s group of G trajectories with rewards Rj , the leave-one-out baseline and group standard deviation are 1 X bi = Rj , si = std{Rj : j ̸= i}, (5) G−1 j̸=i and the trajectory advantage is   Ri − bi Ai = clip , −c, c (si > 0), (6) si + ϵ with a small ϵ and a fixed symmetric clip bound c; zero-variance groups are left unsharpened. The scalar Ai is expanded over the generated tokens of trajectory i. Trajectory grouping uses the original dataset prefix (not the simulator-augmented transcript) so that multi-turn prefixes are grouped correctly. On-policy update with sampling correction. We perform exactly one update per rollout batch and force the policy ratio to one, so the clipped-ratio (PPO/Clip-Higher) term is inactive and no reference KL penalty is applied. A separate token-level truncated-importance-sampling weight corrects the mismatch between training-time and generation-time log probabilities, wi,t = min(τ, exp[log πtrain (yi,t ) − log πgen (yi,t )]) , (7) with a fixed truncation bound τ . With sequence-level aggregation the effective loss is B |yi | 1 X 1 X LGRPO (θ) = − wi,t Ai log πθ (yi,t | xi , yi,