Collaborating with AI Agents: arXiv:2503.18238v3 [cs.CY] 5 Feb 2026 A Field Experiment on Teamwork, Productivity, and Performance Harang Ju Sinan Aral Johns Hopkins Carey Business School MIT Sloan School of Management harang@jhu.edu sinan@mit.edu February 6, 2026 Abstract We examined the mechanisms underlying productivity and performance gains from AI agents using a large-scale experiment on Pairit, a platform we developed to study human-AI collaboration. We randomly assigned 2,234 participants to human-human and human-AI teams that produced 11,024 ads for a think tank. We evaluated the ads using independent human ratings and a field experiment on X which garnered ∼5M impressions. We found human-AI teams produced 50% more ads per worker and higher text quality, while human-human teams produced higher image quality, suggesting a jagged frontier of AI agent capability. Human-AI teams also produced more homogeneous outputs (or diversity collapse) as their ads were more self-similar. The field experiment revealed higher text quality (produced by human-AI teams) improved click-through rates and view-through duration, while higher image quality (produced by human-human teams) improved cost-per-click rates. We found three mechanisms explained these effects. First, human-AI collaboration was significantly more task-oriented, with 25% more task-oriented messages and 18% fewer interpersonal messages. Second, human-AI collaboration displayed more delegation, as participants delegated 17% more work to AI agents than to human partners and performed 62% fewer direct text edits when working with AI. Third, recognition that the collaborator was an AI moderated these effects as participants who correctly identified they were working with AI were more task-oriented and more likely to delegate work. These mechanisms, in turn, explained performance as task-oriented communication improved ad quality, specifically when working with AI, while interpersonal communication reduced ad quality; delegation improved text quality but had no effect on image quality (consistent with the jagged frontier) and was positively associated with diversity collapse, creating homogeneous outputs of higher average quality. The results suggest AI agents drive changes in productivity, performance, and output diversity by reshaping teamwork. 1 1 Introduction Artificial intelligence (AI) tools have garnered attention for their potential to improve productivity and performance (Eloundou et al., 2024; Bick et al., 2024). For example, large language models (LLMs) decreased the average time taken for mid-level professional writing tasks by 40% and increased output quality by 18% (Noy and Zhang, 2023). For job seekers, AI assistance with resumes increased job hiring by an average of 8% (Wiles et al., 2023), and for customer support workers, AI assistance increased productivity by an average of 14% (Brynjolfsson et al., 2023). Moreover, productivity gains were greater for lower-skilled workers (Noy and Zhang, 2023; Brynjolfsson et al., 2023; Choi and Schwarcz, 2023) and varied across task domains (Dell’Acqua et al., 2023). Evidence from an online labor market suggests there has already been a reduction in demand for freelance knowledge work with the advent of generative pre-trained transformer (GPT) models (Hui et al., 2023). In a study reviewing over 106 papers, Vaccaro et al. (2024) showed that human-AI groups outperformed humans alone in 85% of the studies. While studies like Liu et al. (2023) explore levels of proactivity in AI agents, they typically do so without randomized controlled trials (RCTs) measuring productivity effects. The studies that use RCTs to estimate productivity effects tend to randomize access to LLM chatbots [e.g. Dell’Acqua et al. (2023) and Chen and Chan (2024)], which are not typically multimodal, do not include context, do not allow the chatbots to take independent actions or use APIs to call outside of the work environment, and do not provide a collaborative workspace where machines and humans can jointly manipulate output artifacts in real time. These innovations are meaningful because AI agents today possess all these features, yet the existing scientific literature studies none of them. Furthermore, we currently lack fine-grained task level insight into how human-AI collaboration changes work processes and communication patterns and how these changes affect productivity and performance. The vast majority of existing research focuses on the productivity effects of GPT chatbots on individual workers (Noy and Zhang, 2023; Dell’Acqua et al., 2023) or how AI changes people’s perceptions, beliefs, and behaviors (Tey et al., 2024; Costello et al., 2024). However, it is unclear how these interactions evolve in real-time collaborations, especially in 2 environments where AI agents can take autonomous actions, adapt dynamically to human input, and participate in tasks requiring creativity and coordination. Our lack of evidence on in vivo human-AI collaboration exists, in part, because off-the-shelf experimental platforms and studies do not provide collaborative workspaces where researchers can precisely record and measure the collaboration itself: e.g., transcripts of messages between machines and humans, logs of edits to output artifacts, and API (application programming interface) calls to outside agents or tools. In contrast, current AI applications, such as Notion AI and Cursor, already integrate AI agents into such collaborative workspaces and interfaces. To address these gaps, we developed Pairit, a novel experimentation platform designed to study human-AI collaboration in real-world, extensible tasks. Pairit introduces several key innovations. It enables real-time collaboration between humans and AI agents, allowing participants to manipulate text, images, and workflows collaboratively in chat-enabled workspaces that mirror existing online AI collaboration work processes. The platform supports randomized pairings of humans and AI (i.e., human-human or human-AI teams) and allows for randomization of prompts and model fine-tuning. Critically, the AI can perform the complete set of equivalent actions that humans can perform in the collaborative workspace. In the marketing experiment analyzed in this study, these include sending chat messages, writing ad copy, editing ad copy, writing calls to action, editing calls to action, scrolling through images, editing images, selecting images, and generating new images using an external call to Dall-E 3. Moreover, Pairit captures every time-stamped keystroke, message, edit, swipe, scroll, selection, API call, and intermediate output, providing a rich datasets that allow for the detailed reconstruction of collaboration workflows. This represents a fundamental departure from the existing literature, which enables RCT-based evaluation of the outputs and productivity implications of human collaboration with LLM-based chatbots and co-pilots, but does not enable randomized experiments analyzing the task-level productivity and work process changes created in human collaborations with fully functioning, multimodal AI agents. To study the work process, productivity, and performance implications of human collaboration with multi-modal AI agents, we conducted a large-scale randomized study of human-AI collabo- 3 ration on advertising design and creation, a task requiring creativity, iteration, and precision. A total of 2,234 participants, representative of the U.S. population, were recruited through Prolific and randomly assigned to human-human or human-AI teams using the Pairit platform. Teams worked collaboratively to create marketing campaigns for a think tank’s year-end annual report, including generating and selecting ad images and writing ad copy and calls to action. This process was fully recorded, resulting in a dataset including 11,024 ads, 182,607 messages, 1,889,559 text edits, 62,119 image edits, and 10,074 AI-generated images, offering an unprecedented level of detail with which to understand work processes, communication, productivity and output quality. Once the lab portion of the experiment was completed, we conducted a field experiment on the ads produced by human-human and human-AI teams. We obtained human and AI quality ratings of the ads, including the quality of the ad copy and images, and the human- and AI-evaluated likelihood of consumer engagement with the ads (measured by click-through rates). We then ran the ads in a real online display ad ecosystem, generating over 4.9 million impressions on X, and evaluated click-through rate, cost-per-click, view-through rate, and view-through duration metrics on the annual report using the platform’s ads API and DocSend’s view metrics, which allow us to record how much of the report consumers read, page by page, after clicking through on the ads. Our analyses examine three performance effects documented in the AI literature but in an agentic collaboration context: productivity and performance (Noy and Zhang, 2023; Brynjolfsson et al., 2023), the jagged frontier of AI capabilities across tasks (Dell’Acqua et al., 2023; Gans, 2026), and diversity collapse, or the homogenization of AI-assisted outputs (Padmakumar and He, 2024; Chen and Chan, 2024; Hao et al., 2026). We found evidence of all three effects in our agentic setting: human-AI teams produced 50% more ads per worker, higher-quality text, and more homogeneous outputs, while human-human teams produced higher-quality images. Field tests revealed that higher image quality (from human-human teams) improved cost-per-click while higher text quality (from human-AI teams) improved click-through rates and view-through rates, with these offsetting quality effects explaining similar overall field experiment ad performance across team types. Most importantly, we found that specific teamwork dynamics—task-orientation, delegation, 4 and AI recognition—moderate these effects. While our treatment randomizes AI agents at the extensive margin (humans working with or without AI), the data we collected on communication patterns, delegation decisions, and AI recognition capture how people interact with, collaborate with, and delegate work to AI agents. In mining the rich collaboration data produced by Pairit, we found individuals in human-AI teams sent 62% more messages than those in human-human teams. Furthermore, human-AI teams sent 25% more content- and process-oriented messages, especially messages containing suggestions, instructions, prioritization, and planning. Conversely, human-human teams sent 18% more social and emotional messages, including messages that expressed rapport building, self-assessment, and concern. Task-orientation, when working with AI, was then associated with higher text quality, image quality and likelihood of clicking, while interpersonal communication was associated with lower text quality, image quality and likelihood of clicking. We also tracked work process changes: human-AI teams made 62% fewer direct edits to the copy and delegated 17% more work to their AI partners. Higher delegation was then associated with improved text quality but not image quality, and with reduced output diversity. To support the idea that AI reshapes teamwork, we found that participants who correctly recognized that they were working with AI were more task-oriented and more likely to delegate work to their AI partners, teamwork processes associated with higher performance. Together, these results strongly suggest that collaboration with AI agents reshapes work processes in specific ways and that these new teamwork dynamics moderate productivity gains and jagged frontier quality effects in human-AI collaboration. Although prior studies have shown that AI tools can improve productivity and reduce task completion times (Noy and Zhang, 2023; Brynjolfsson et al., 2023), AI is often treated as a passive tool rather than as an active collaborator in prior research. As AI agents become integral to modern workflows, researchers are beginning to explore their role as work collaborators, rather than mere tools, emphasizing the importance of trust, transparency, and integration in human-AI partnerships (Makarius et al., 2020; Anthony et al., 2023; Collins et al., 2024). Our work contributes to this emerging literature by presenting the first task-level randomized experiment measuring the work 5 process, productivity and performance implications of human-AI collaboration with fully functional, multimodal AI agents. The Pairit platform serves as the methodological infrastructure enabling this comparison, and the field experiment provides external validation of quality and performance differences. The core contribution is the empirical analysis of how human-AI collaboration reshapes task-orientation, delegation, communication, and recognition, as well as output quality, performance, and productivity. We hope these observations inform future research as AI agents become integral to the workplace. 2 Theory 2.1 Teamwork in Human-Human and Human-AI Teams Effective teamwork requires more than individual competence and depends critically on how team members coordinate, communicate, and manage interpersonal relationships (Marks et al., 2001). These social dynamics generate collective intelligence that exceeds the sum on individual abilities (Woolley et al., 2010). A foundational distinction in this literature separates taskwork, the technical work itself, from teamwork, the processes that enable individuals to work together effectively (Salas et al., 1992). Teamwork encompasses three types of processes: transition processes where teams plan and set goals, action processes where they coordinate and monitor progress, and interpersonal processes where they manage conflict and build motivation (Marks et al., 2001). Teams cycle through these phases as they execute work. Human teams invest substantial effort in all three. They build shared mental models through ongoing communication about goals, strategies, and progress, enabling them to anticipate each other’s actions (Mathieu et al., 2000). This shared understanding allows teams to negotiate who does what and when. Members also build trust, provide support, and resolve conflict to maintain cohesion, though this consumes time and cognitive resources. The emergence of AI agents that can act autonomously within collaborative workspaces represents a fundamental shift from AI as a tool to AI as a teammate. Unlike chatbot-style interfaces 6