TVS Sidekick: Challenges and Practical Insights from Deploying Large Language Models in the Enterprise Paula Reyero Lobo1,2 , Kevin Johnson1 , Bill Buchanan1 , Matthew Shardlow2 , Ashley Williams2 , Samuel Attwood2 1 TVS Supply Chain Solutions, Chorley, UK 2 Manchester Metropolitan University, Manchester, UK {P.ReyeroLobo, M.Shardlow, Ashley.Williams, S.Attwood}@mmu.ac.uk {kevin.johnson, bill.buchanan}@tvsscs.com Abstract Many enterprises are increasingly adopting Ar- arXiv:2509.26482v1 [cs.AI] 30 Sep 2025 tificial Intelligence (AI) to make internal pro- cesses more competitive and efficient. In re- sponse to public concern and new regulations for the ethical and responsible use of AI, im- plementing AI governance frameworks could Figure 1: Overview of TVS Sidekick, an AI assistant help to integrate AI within organisations and that leverages LLMs to answer queries with relevant mitigate associated risks. However, the rapid enterprise data using retrieval augmented generation technological advances and lack of shared eth- (RAG) via a Microsoft Teams extension. ical AI infrastructures creates barriers to their practical adoption in businesses. This paper presents a real-world AI application at TVS efficiency. TVS SCS UK has decided not to use Supply Chain Solutions, reporting on the expe- third-party software integrators or product vendors rience developing an AI assistant underpinned for its solutions, which would negatively impact by large language models and the ethical, regu- their agility and innovation. Instead, they have latory, and sociotechnical challenges in deploy- started their journey towards an AI transformation ment for enterprise use. through an in-house AI team. TVS Sidekick is the flagship product of this in- 1 Introduction house team. TVS Sidekick is built upon the prin- Recent developments are driving industry interest ciples of Retrieval Augmented Generation (RAG) in the field of Large Language Models (LLMs). (Lewis et al., 2020). All relevant company docu- Key developments of note are the abundant avail- ments, as available via their internal cloud-based ability of commercial language modelling solutions systems, are vectorised and compared to the input (Devlin et al., 2019; Brown et al., 2020; Thoppilan query, with the LLM then performing information et al., 2022) and the increased public awareness of extraction for the purposes of question answering the capabilities of LLMs (Mialon et al., 2023; Qu with custom prompting (Qu et al., 2025). Users et al., 2025). However, to successfully utilise these interact with TVS Sidekick via a Microsoft Teams models, organisations must navigate important so- extension (Figure 1). cietal challenges related to ethics, sustainability, As TVS SCS UK advances its AI transformation and compliance (Hagendorff, 2024; Laux et al., through the development of Sidekick, it must also 2024). navigate a complex legal and regulatory landscape. TVS SCS UK is a top-tier third-party logistics At the centre of this landscape are the European (3PL) provider in Europe and the UK, offering com- Union Artificial Intelligence Act (EU AIA) (Eu- prehensive supply chain solutions. 3PL customers ropean Commission, 2014) and related standards, increasingly adopt intelligent technology-led so- such as ISO/IEC 42001 for AI Management Sys- lutions to optimise their supply chain operations tems (International Organization for Standardiza- and reduce costs (Pournader et al., 2021; Li et al., tion., 2023). Furthermore, TVS SCS UK must 2023). To stay ahead of the competition, TVS SCS overcome a range of sociotechnical challenges that UK are leveraging LLMs to create a competitive accompany the deployment of LLMs, such as is- advantage and enhance their internal operational sues of fairness, transparency, and accountability (Crockett et al., 2023; Ojewale et al., 2025), which superiority in fine-tuning applications (Devlin et al., limit their practical adoption in enterprise environ- 2019). With increased data size and model com- ments. plexity, decoder-only models like the generative In this paper, we report on TVS SCS UK’s expe- pre-trained transformer (GPT) model series have rience developing Sidekick, navigating the relevant become more attractive for industry due to their legislation and regulations, and overcoming the few/zero-shot performance (Brown et al., 2020). challenges they have encountered along the way. This paradigm shift led to methods for aligning to user intent (Ouyang et al., 2022) (like reinforce- 1.1 Significance of this Study ment learning with human feedback) powering pop- This study presents practical insights from applied ular conversation-focused products like ChatGPT. AI research in a real-world business context. To be While these scaled-up models offer business value specific, we contribute to the field in three ways: (e.g. analysing vast data in real-time), issues such as the closed-source nature of existing solutions • Technical Contributions. We describe the de- creates barriers to organisations lacking computa- sign and implementation of Sidekick, an AI tional power (Yang et al., 2024). assistant underpinned by LLMs that is tailored Focus on approaches including RAG (and for enterprise use, including novel approaches pipeline parts showing improvement). Recently, to prompt engineering and RAG. the focus has turned into giving more agency to • Regulatory Contributions. We present a case LLMs to become independent problem solvers. For study of how a business is aligning its devel- instance, by consulting with external knowledge opment with emerging legislation and regula- sources for factual grounding (Lewis et al., 2020; tions, most notably the EU AIA, by working Thoppilan et al., 2022). More broadly, a significant towards harmonised technical standards (e.g., step forward is the combination of “tools”, namely ISO/IEC 42001). tool-augmented LLMs (Mialon et al., 2023), in- cluding retrieval-augmented language models for • Sociotechnical Contributions. We explore the efficiently handling new data. Such approaches sociotechnical challenges that accompany the generally consist of four stages: task planning (i.e. deployment of LLMs in enterprise environ- break down user query into tasks), tool selection, ments. We report quantitative statistics relat- tool calling, and response generation (Qu et al., ing to the adoption of Sidekick alongside a 2025). Similarly, critical advances require frame- qualitative analysis of end-user feedback. works for enabling LLMs to recall previous interac- tions (Zhang et al., 2024), allowing for multimodal 1.2 Structure of this Study data processing (Sun et al., 2025; Song et al., 2025), The remainder of this paper is organized as follows. or to improve responses based on past interactions Section 2 reviews the relevant literature. Section (Wang et al., 2024). 3 details the technical implementation of Sidekick. This paper presents a case study of recent LLM Section 4 presents a case study of how TVS SCS developments in practice, specifically through the UK is aligning its development with emerging leg- technical implementation of an AI assistant that islation and regulations. Section 5 includes a quan- processes heterogeneous enterprise data sources us- titative and qualitative evaluation of the progress ing knowledge augmentation strategies, including to date. Finally, section 6 concludes this paper and novel approaches to prompt engineering and RAG. describes directions for future work. 2.2 Responsible and Ethical AI 2 Related Work Challenges in training, evaluating, and deploy- 2.1 LLMs in the Enterprise ing LLMs and emerging AI regulation. While Prominent applications, reviewing strategies to AI shows great potential and business opportuni- augment LLM capabilities. The transformer archi- ties, many concerns arise from embedding biases, tecture enhanced language modelling capabilities contributing to climate degradation, threatening hu- and has since sparked great attention in industry man rights and more (UNESCO, 2021). An active (Vaswani et al., 2017). This led to many readily research area has emerged for responding hard nor- available pre-trained models, which proved their mative questions related to AI, such as bias and Principles Requirements Human oversight & accountability AI to support/augment humans, with humans clearly accountable. Technical robustness and safety AI tools work as expected, minimising potential harms. Transparency Clear notification of AI involvement, clear and traceable outputs. Privacy & data governance Follow existing privacy rules with quality, robust data. Diversity & fairness Output free of bias and does not discriminate or treat unfairly. Social & environmental wellbeing AI is sustainable and beneficial to all. Table 1: Key emerging principles and requirements from global AI regulations (British Standards Institution, 2025). fairness, transparency, and accountability (Jobin From theory to practice. While approaches to et al., 2019). Institutions at global, international, ethical AI exist (including bias tests, checklists and and national levels have responded with recommen- risk impact assessments), organisations face bar- dations for responsible and ethical AI, consisting riers that limit their practical adoption (Crockett of principles and practices such as a human rights- et al., 2023). Technical approaches alone are not centred approach to AI (UNESCO, 2021), or AI sufficient to establish an ethical AI infrastructure assurance methodologies (i.e. to “measure, eval- (Ojewale et al., 2025). Instead, participatory ap- uate, and communicate the trustworthiness of AI proaches involving civil society stakeholders are systems” (Department for Science, Innovation & needed for effective standard setting, implementa- Technology, 2024)). The advent of LLMs only tion, and enforcement (Crockett et al., 2024; Mod- adds a layer of complexity to the ethical debate hvadia et al., 2025). This paper contributes to bridg- (Hagendorff, 2024), raising additional concerns (re- ing the gap between theory and practice through the garding transparency, copyright, and safety) (Eu- experience of implementing an AI governance strat- ropean Commission, 2025) that require specific egy in a real-world business context, reporting on regulation for generative AI technologies. the technical, legal and human challenges involved Global legislation and EU AIA as most far with the adoption of generative AI technologies. reaching and punitive of regulations. The EU AIA 2.3 Positioning this Study is a notable example leading the field of AI regu- In the logistics sector, real-time data analysis can lation, with significant non-compliance penalties transform business operations, from internal ware- to business providing or deploying AI. While leg- housing and inventory processes to stakeholder islation approaches and requirements vary across management (Pournader et al., 2021). However, jurisdiction areas (Table 1), AI regulations are de- empirical research in related areas (Qian et al., veloping globally to provide assurances in critical 2024; Kapania et al., 2025) shows that benefits and aspects such as human oversight and accountability, trade-offs in the use of AI technologies manifest technical robustness and safety, or privacy and data differently depending on their application domain. governance (British Standards Institution, 2025). Despite growing understanding of public at- Governments and legislative bodies are working titudes towards AI (Modhvadia et al., 2025; towards practical strategies to implement the prin- Mhasakar et al., 2025), research on its industrial ciples underlying AI regulations. Harmonised stan- application remains limited. This study presents dards are one of the primary mechanisms for help- insights from the development and use of LLMs at ing organisations translate regulatory requirements TVS SCS UK, to address the following gaps: into technical implementations (AI Standards Hub, • Examining the implementation and practical 2024). Standardisation should specify minimum application of recent LLM advances within technical testing, documentation, and public report- the enterprise context. ing to limit AI developers and/or users discretion in complying with regulatory requirements (Laux • Embedding high-level ethical principles in et al., 2024). However, local empirical studies AI regulatory frameworks into organisational and specific examples of how organisations imple- practices. ment processes that ensure AI regulation principles • Empirical analysis of challenges that emerge (Wolf-Brenner et al., 2024) is crucial for a demo- with adopting LLMs in a logistics company. cratic approach to ethical and responsible AI. Figure 2: Architecture diagram showing the main components of Sidekick, namely the ingestion and RAG pipelines, with novel approaches to prompt engineering (to handle code queries) and augmentation retrieval (for tool use). 3 Technical Implementation • Indexing. Extracting semantic vectors from each chunk with an embedding model, and This section presents the design and implemen- creating an index in the vector database for tation of Sidekick (Figure 2), describing: (i) the each data source (to define specific fields). integration of relevant company data into a vector database (Ingestion pipeline), and (ii) how this vec- Sidekick is developed to handle both text and torised data is used to process user queries with code-related queries. Crucially, using prompt engi- enhanced LLM capabilities (RAG pipeline). neering for code integration. First, files are split to 3.1 Ingestion Pipeline objects by logical meaning (i.e. functions, methods, or procedures). An LLM is prompted to generate Vector databases are increasingly used to enhance descriptions to each code file, using its object list LLM-generated outputs by providing relevant text to report on the overall purpose, structure, key pro- fragments (“chunks”) that have a similar mean- cedures, functions, and external interactions. Both ing to the user query (i.e. “context”). To do so, code and transformed text fragments are stored in company data needs to be transformed and embed- the vector database, to expose relevant source code ded into a common database that handles semantic lines as sources when responding to the user query. similarity searches. The vector database acts as a bridge between the two system components, accel- 3.2 RAG Pipeline erating the retrieval of content that is relevant to the user query. The second system component processes user The first system component integrates informa- queries by leveraging company data and conver- tion from different company data sources into the sation history to enhance LLM outputs. vector database, in two main steps: The user query and conversation history (i.e. queries and responses of the last 60-minute ses- • Data preparation. First, TVS data is fetched sion) are sent to a router. The router splits the user from different data sources, i.e. SharePoint, query into sub-sentences (i.e. specific tasks) and Azure DevOps (ADO), code repositories, and calls an LLM to decide which route to take for the TVS website, with a scheduled hour refresh. augmentation retrieval. Each route uses a type of Data is then processed to extract chunks using “chatterbot”, a tool-based LLM optimised to answer a document loader: i.e. parsing (extract or questions related to different data sources. transform to text - for code) and chunking Each task identified from the user query triggers (splitting by semantic or logical boundaries). an instance of RAG: Requests Standard(s) 42001 Requirement Focus Accuracy 23282* 4.[1/2/3] Purpose & Requirements Robustness 24027, 12791 6.[2/3] Objectives & Change Transparency 12792* 5.[1/2] Leadership & Policy Human oversight 8200, 42105* 5.3 Roles & responsibilities Data and Data Management 25012, 5259 6.1.[1/2/3], AI Risks Y Cybersecurity 27001 8.[1/2/3/4] Record keeping and logging 24970* 9.1. 9.2.[1/2] Monitoring & Measuring Y Quality management systems 9001, 25059* 10.[1/2], 9.3 Continuous improvement Y Risk management systems 31000, 23894 7.[1/2/3/4], Awareness & Training Conformity assessment 42006 7.5.[1/2/3] Table 2: Horizontal standardisation request for the EU Table 3: Mapping analysis between ISO/IEC 42001 AIA (AI Standards Hub, 2024), mapped to available and existing management systems at TVS SCS UK, ISO/IEC standards. Highlighted standards (*) are yet to highlighting focus areas for implementation (“Y”). be published (20th August 2025). Different harmonised standards are being devel- • Retrieval: information retrieval from vector oped to support the implementation of the EU AIA, database using the same embedding model to such as the ISO/IEC 12792 and 24970 standards for extract context (top-10 similar chunks) and re- addressing the transparency and logging of AI sys- format chunks (its text and metadata as XML tems, respectively (see Table 2). Building upon rel- or JSON list for code route). evant standards, including AI Concepts and Termi- nology (22989) and AI Risk Management (23894), • Augmentation: calls an LLM to extract the ISO/IEC 42001 is the first international standard required parameters to generate the answer for AI Management Systems, aiming to guide or- (including prompt template). ganisations in the responsible development and use • Generation: calls an LLM using the instruc- of AI systems. tions and context from previous steps. Recognising the value of standards to opera- tionalise AI regulation principles for ethical and The output generated for each task are combined responsible AI, TVS SCS UK has decided to adopt into a single response using the LLM only with gen- an AI Management System (AIMS) framework to erated texts. The user query and response are saved develop trustworthy AI solutions. for logging and leveraging conversation history. 4.2 ISO/IEC 42001 Implementation 4 Navigating Regulatory Challenges of TVS SCS UK have developed and deployed for- TVS Sidekick: Case Study mal management systems in important areas such This section presents the regulatory challenges that as information security, quality, health and safety, emerge with the development of LLMs, and how business continuity, and environmental manage- they may be overcome in a real-world business ment. To effectively implement an AI management context. Specifically, we present a case study on system, TVS SCS UK began with mapping the key navigating a complex and changing AI regulatory requirements of ISO/IEC 42001 to existing stan- landscape in the enterprise, leading to the imple- dards, focusing on management systems already mentation of the first harmonised technical stan- adopted by the organisation. dard for responsible AI development and use. The results from this mapping analysis are shown in Table 3. Notably, TVS SCS UK maintains 4.1 EU AIA & Harmonised Standards an Information Security management system fol- TVS SCS UK is achieving compliance working lowing ISO/IEC 27001 (International Organization towards AI standardisation, which is key to the for Standardization., 2022). Processes supporting development and adoption of AI. One key regula- this standard, especially related to data manage- tion shaping the field of standardisation is the EU ment and cybersecurity, were aligned with ISO/IEC AIA, which is leading the global landscape of AI 42001 requirements. This comparison helped to regulation. identify focus areas for developing an AIMS: Category Topics Usage indicators Performance alignment, reliability, robustness, Interaction volume Number of messages (i.e. prompt engineering, usefulness, prompts) and unique users. helpfulness, truthfulness Response time Average response time (s). Safety privacy, security, safety User engagement Average of messages per interpretability, transparency, session (on daily basis). explainability, fairness, trust- worthiness, adversarial attacks Table 5: Description of metrics in the monitoring system supporting the AI assitant at TVS SCS UK. Regulation regulation, best practice*, gover- nance, compliance, accountability Continuous improvement. The effective man- Table 4: Topics of AI/LLM performance, safety, and regulation feeding into the Knowledge Base of papers. agement of vulnerabilities to the AIMS is crucial for demonstrating continual improvement in the use of AI, with documented validation and verifi- cation. TVS SCS UK is establishing processes for AI Risks. TVS SCS UK maintains a risk man- maintaining and deploying AI, primarily focused agement strategy as an integral part of their infor- on the evaluation and technical documentation of mation security. This ongoing process sets out Sidekick. To this end, a primary evaluation objec- responsibilities and a methodology to periodically tive has been set to understand the needs and ways assess risks based on likelihood and impact levels. in which the AI assistant may best support differ- One of the main challenges introducing AI is the ent company roles and responsibilities. Specifi- need of staying relevant with current risks. To this cally, through the organisation of periodic feedback end, TVS SCS UK is working towards establishing interviews as part of a continuous evaluation of a Knowledge Base that informs AI development Sidekick, with target populations whose adoption and use within the company. The AI team started of AI could bring most benefit to the company. A maintaining an academic database of research re- participatory approach to AI development aims to views including meta-analyses and relevant case support a culture of ethical and responsible AI. studies that is accessible throughout the company; both in full-text and via their in-house AI assistant 5 Monitoring & Evaluation for the purpose of question answering. Further- more, a systematic search (Brereton et al., 2007) This section presents insights gathered from the de- of AI research papers in relevant topics (Table 4) ployment of LLMs at TVS SCS UK, highlighting allows to explore topic distribution and relevant sociotechnical challenges in their enterprise use. metadata, such as indexed keywords or keywords Following the on-going implementation of an AI from the authors, and supports the maintenance and governance model, we specifically report on empir- updating of the academic database. ical findings from the monitoring system and initial Monitoring & Measuring. Another component evaluation of the Sidekick product. of the AIMS framework is to capture monitors and 5.1 Adoption & Usage measures on the use of AI, including an internal au- dit programme. TVS SCS UK is developing a mon- The implementation of an AIMS framework follow- itoring system supporting Sidekick, which includes ing ISO/IEC 42001, in particular related to Moni- usage indicators (Table 5) and descriptive metrics toring & Measuring requirements, provides prac- of interactions (volume breakdown by department, tical insights on the levels of AI adoption and us- job title, individual user, and question type). Ul- age in the organisation. Consequently, we report timately, these metrics aim to pragmatically mea- findings from the monitoring system described in sure the effectiveness of the AI assistant, setting a Section 4.2. starting point for other AI performance and safety Figure 3 shows quantitative statistics related to measures. For instance, obtained through the provi- the initial adoption of Sidekick at TVS SCS UK. sion of feedback channels (Torkamaan et al., 2024) The monitoring system shows usage indicators and to report quality or safety incidents, or the inclu- descriptive metrics of interaction volume within a sion of LLM observability evaluations (Kenthapadi 4-month period (March-June 2025). et al., 2024). Overall, continued use of the AI assistant is Figure 3: Monitoring system measuring real-time usage data of TVS Sidekick. shown within the observed time period. This is roles correspond to IT (business technology, busi- seen in continued measures both in terms of user ness development), management, and operational engagement and interaction volume (exceeding 500 areas such as defence, technology, commercial, and prompts in the first two months and 250 in the fol- operations. lowing two months). Despite fluctuations, the con- In terms of individual usage, the breakdown of versations do not seem to be long, rarely exceeding activity per individual makes a clear distinction an average of five questions per conversation. The between lead and early adopters (i.e. 46 - 253 response time has peaks on specific dates that in- queries) and occasional users (less than 20 on av- crement the time to 46 seconds on average. erage). Queries answered with SharePoint data The descriptive analysis of interaction volume (i.e. question) were the most common, followed at organisational level reveals that the most active by responses without retrieval augmentation (gen- users were primarily in technology (e.g. devel- eral response), codebase file queries (rpg query) opers) and business roles (e.g. bid management, and queries related the development environment, business analysts). At the departmental level, these i.e. Azure DevOps (ado query). Understanding current use of AI/Sidekick guage with limited technical documentation. The Have you used Sidekick/other AI tools? current version of the AI assistant offers a good What have you used it for? starting point for understanding key parameters Where was AI/Sidekick most helpful/ and functions, but remains limited in addressing unhelpful? more specific queries from technical users. Outlook Keen to engage with AI (Mentioned by: 11). In what aspects of your job would AI Overall, staff were enthusiastic about using Side- be most useful? kick to standardise code, reduce duplication, re- Do you have any concerns about inte- fer new starters to source documentation, or avoid grating AI into your workflow? ownership issues when using external AI tools. Fur- thermore, new features were proposed, including Table 6: Topic guide of feedback interviews supporting learning from user prompts or returning questions the continuous evaluation of TVS Sidekick. to users to resolve ambiguous queries. Privacy/commercially sensitive questions/Other 5.2 Qualitative Feedback concerns (Mentioned by: 8). There were no major concerns with the use of Sidekick, provided it was The initial round of feedback interviews that feed fed with the right information and access levels. into the Continuous improvement requirement un- Concerns were raised around job security and dis- der ISO/IEC 42001 highlights significant chal- trust in AI tools, along with the emphasis on using lenges when introducing AI in the business context. Sidekick internally due to potential disclosure of Primarily, with respect to the perceived benefits information from the client side. and risks of deploying LLMs in the enterprise, due The first round of feedback has been worked into to Sidekick being the flagship product. a plan for continual improvement and addressing In total, 24 interviews with members of the IT concerns, informing further developments of TVS department at TVS SCS UK were conducted be- Sidekick. TVS SCS UK will continue developing tween March and April 2025. Participants were processes to adhere to ethical principles in regula- invited to 30-minute online meetings for a semi- tory standards, sharing practical insights in critical structured interview. The topic guide (Table 6). areas such as managing AI risks, providing relevant included questions i) to gather experiences so far monitors and measures on AI use, and increasing in using AI/Sidekick at work and ii) understand AI adoption through training and consultation. how TVS staff want to use Sidekick in the future. Finally, interview minutes were thematically anal- 6 Conclusion ysed (Byrne, 2022) by two independent coders. The analysis of qualitative feedback led to better This paper presented the experience and challenges understanding of baseline attitudes towards AI. The encountered in a real-world business scenario with following themes were identified: the development and deployment of TVS Sidekick, Enhanced retrieval (Mentioned by: 16). A key an AI assistant leveraging LLMs for enterprise use. advantage of Sidekick over other tools is its speci- This empirical study provides practical knowledge, ficity to TVS data. Users valued its assistance with including key lessons learned from the implemen- SharePoint-related tasks, finding it faster than a tation and governance of the in-house AI assistant. manual search and with a “readable and visible” Limitations & Ethical Considerations format, especially for the source list. Good extracting business logic (Mentioned by: The findings and insights presented are drawn from 10). Sidekick was particularly helpful in providing a specific organisational context and reflect expe- business knowledge, with clear use cases for busi- riences within a particular time frame and initial ness analysts. Specifically, for understanding the phase of evaluation. While the technical specifics context of TVS data and key definitions of compo- and detailed implementation of each component of nents within business processes. the governance framework are outside the scope of Not enough technical detail (Mentioned by: 13). this work, this paper aims to contribute to the wider Developers emphasized the need for more domain community by sharing reflections on navigating knowledge to explain internal programmes. Partic- technical, ethical, regulatory, and sociotechnical ularly, those relying on a legacy programming lan- challenges of deploying LLMs in practice. References North American chapter of the association for com- putational linguistics: human language technologies, AI Standards Hub. 2024. Demystifying the EU AI volume 1 (long and short papers), pages 4171–4186. Act – Implications for UK businesses. Available at: https://aistandardshub.org/events/inno European Commission. 2014. Regulation (EU) vate-uk-bridgeai-demystifying-the-eu-a 2024/1689 of the European Parliament and of the i-act-implications-for-uk-businesses/ Council of 13 June 2024 (Artificial Intelligence Act). [Accessed: 20th August 2025]. Available at: https://eur-lex.europa.eu/e li/reg/2024/1689/oj/eng [Accessed: 20th Au- Pearl Brereton, Barbara A. Kitchenham, David Budgen, gust 2025]. Mark Turner, and Mohamed Khalil. 2007. Lessons from applying the systematic literature review pro- European Commission. 2025. The General-Purpose AI cess within the software engineering domain. Jour- Code of Practice. Available at: https://digita nal of Systems and Software, 80(4):571–583. Soft- l-strategy.ec.europa.eu/en/policies/c ware Performance. ontents-code-gpai [Accessed: 20 August 2025]. British Standards Institution. 2025. Understanding Thilo Hagendorff. 2024. Mapping the ethics of genera- ISO/IEC 42001 – a framework for managing AI. tive ai: A comprehensive scoping review. Minds and Available at: https://iuk-business-conne Machines, 34(4):39. ct.org.uk/events/understanding-iso-i International Organization for Standardization. 2022. ec-42001-a-framework-for-managing-ai/ Information security, cybersecurity and privacy pro- [Accessed: 20th August 2025]. tection — Information security management systems — Requirements (ISO Standard No. 27001:2022). Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Available at: https://www.iso.org/standard Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind /27001 [Accessed: 20th August 2025]. Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, International Organization for Standardization. Gretchen Krueger, Tom Henighan, Rewon Child, 2023. Information technology —- Artificial Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Intelligence – Management Systems (ISO Clemens Winter, Christopher Hesse, Mark Chen, Eric Standard No. 42001:2023). Available at: Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, https://www.iso.org/standard/42001 Jack Clark, Christopher Berner, Sam McCandlish, [Accessed: 20th August 2025]. Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Anna Jobin, Marcello Ienca, and Effy Vayena. 2019. Proceedings of the 34th International Conference on The global landscape of ai ethics guidelines. Nature Neural Information Processing Systems, NIPS ’20, machine intelligence, 1(9):389–399. Red Hook, NY, USA. Curran Associates Inc. Shivani Kapania, Ruiyi Wang, Toby Jia-Jun Li, Tianshi David Byrne. 2022. A worked example of Braun Li, and Hong Shen. 2025. ’i’m categorizing llm as and Clarke’s approach to reflexive thematic analy- a productivity tool’: Examining ethics of llm use in sis. Quality & Quantity, 56(3):1391–1412. hci research practices. Proc. ACM Hum.-Comput. Interact., 9(2). Keeley Crockett, Edwin Colyer, Lauren Coulman, Krishnaram Kenthapadi, Mehrnoosh Sameki, and Ankur Caitlin Nunn, and Sarah Linn. 2024. Peas in pods: Taly. 2024. Grounding and evaluation for large Co-production of community based public engage- language models: Practical challenges and lessons ment for data and ai research. In 2024 International learned (survey). In Proceedings of the 30th ACM Joint Conference on Neural Networks (IJCNN), pages SIGKDD Conference on Knowledge Discovery and 1–10. Data Mining, KDD ’24, page 6523–6533, New York, NY, USA. Association for Computing Machinery. Keeley Crockett, Edwin Colyer, Luciano Gerber, and Annabel Latham. 2023. Building trustworthy ai so- Johann Laux, Sandra Wachter, and Brent Mittelstadt. lutions: A case for practical solutions for small busi- 2024. Three pathways for standardisation and ethi- nesses. IEEE Transactions on Artificial Intelligence, cal disclosure by default under the european union 4(4):778–791. artificial intelligence act. Computer Law & Security Review, 53:105957. Department for Science, Innovation & Technology. 2024. Introduction to AI assurance. Available at: Patrick Lewis, Ethan Perez, Aleksandra Piktus, https://www.gov.uk/government/publicat Fabio Petroni, Vladimir Karpukhin, Naman Goyal, ions/introduction-to-ai-assurance [Ac- Heinrich Küttler, Mike Lewis, Wen-tau Yih, cessed: 20th August 2025]. Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-augmented generation for Jacob Devlin, Ming-Wei Chang, Kenton Lee, and knowledge-intensive nlp tasks. In Proceedings of Kristina Toutanova. 2019. Bert: Pre-training of deep the 34th International Conference on Neural Infor- bidirectional transformers for language understand- mation Processing Systems, NIPS ’20, Red Hook, ing. In Proceedings of the 2019 conference of the NY, USA. Curran Associates Inc. Beibin Li, Konstantina Mellou, Bo Zhang, Jeevan Jiankai Sun, Chuanyang Zheng, Enze Xie, Zhengying Pathuri, and Ishai Menache. 2023. Large Language Liu, Ruihang Chu, Jianing Qiu, Jiaqi Xu, Mingyu Models for Supply Chain Optimization. arXiv e- Ding, Hongyang Li, Mengzhe Geng, Yue Wu, Wen- prints, page arXiv:2307.03875. hai Wang, Junsong Chen, Zhangyue Yin, Xiaozhe Ren, Jie Fu, Junxian He, Yuan Wu, Qi Liu, Xihui Manas Mhasakar, Rachel Baker-Ramos, Benjamin Liu, Yu Li, Hao Dong, Yu Cheng, Ming Zhang, Carter, Evyn-Bree Helekahi-Kaiwi, and Josiah Hes- Pheng Ann Heng, Jifeng Dai, Ping Luo, Jingdong ter. 2025. ”i would never trust anything western”: Wang, Ji-Rong Wen, Xipeng Qiu, Yike Guo, Hui Kumu (educator) perspectives on use of llms for cul- Xiong, Qun Liu, and Zhenguo Li. 2025. A sur- turally revitalizing cs education in hawaiian schools. vey of reasoning with foundation models: Concepts, In Proceedings of the Extended Abstracts of the CHI methodologies, and outlook. ACM Comput. Surv., Conference on Human Factors in Computing Systems, 57(11). CHI EA ’25, New York, NY, USA. Association for Computing Machinery. Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Grégoire Mialon, Roberto Dessı̀, Maria Lomeli, Christo- Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, et al. foros Nalmpantis, Ram Pasunuru, Roberta Raileanu, 2022. Lamda: Language models for dialog applica- Baptiste Rozière, Timo Schick, Jane Dwivedi-Yu, tions. arXiv preprint arXiv:2201.08239. Asli Celikyilmaz, et al. 2023. Augmented language models: a survey. arXiv preprint arXiv:2302.07842. Helma Torkamaan, Steffen Steinert, Maria Soledad Pera, Olya Kudina, Samuel Kernan Freire, Himan- Roshni Modhvadia, Tvesha Sippy, Octavia Field Reid, shu Verma, Sage Kelly, Marie-Therese Sekwenz, Jie and Helen Margetts. 2025. How do people feel about Yang, Karolien van Nunen, Martijn Warnier, Frances ai? (Ada Lovelace Institute and The Alan Turing Brazier, and Oscar Oviedo-Trespalacios. 2024. Chal- Institute) https://attitudestoai.uk/. lenges and future directions for integration of large language models into socio-technical systems. Be- Victor Ojewale, Ryan Steed, Briana Vecchione, Abeba haviour & Information Technology, 0(0):1–20. Birhane, and Inioluwa Deborah Raji. 2025. Towards ai accountability infrastructure: Gaps and opportuni- UNESCO. 2021. Ethics of Artificial Intelligence | UN- ties in ai audit tooling. In Proceedings of the 2025 ESCO. Available at: https://www.unesco.org CHI Conference on Human Factors in Computing /en/artificial-intelligence/recommenda Systems, CHI ’25, New York, NY, USA. Association tion-ethics [Accessed: 20th August 2025]. for Computing Machinery. Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Carroll Wainwright, Pamela Mishkin, Chong Zhang, Kaiser, and Illia Polosukhin. 2017. Attention is all Sandhini Agarwal, Katarina Slama, Alex Ray, et al. you need. Advances in neural information processing 2022. Training language models to follow instruc- systems, 30. tions with human feedback. Advances in neural in- formation processing systems, 35:27730–27744. Siyuan Wang, Zhongyu Wei, Yejin Choi, and Xiang Ren. 2024. Symbolic working memory enhances language Mehrdokht Pournader, Hadi Ghaderi, Amir Hassan- models for complex rule application. In Proceedings zadegan, and Behnam Fahimnia. 2021. Artificial of the 2024 Conference on Empirical Methods in intelligence applications in supply chain manage- Natural Language Processing, pages 17583–17604, ment. International Journal of Production Eco- Miami, Florida, USA. Association for Computational nomics, 241:108250. Linguistics. Crystal Qian, Emily Reif, and Minsuk Kahng. 2024. Christof Wolf-Brenner, Viktoria Pammer-Schindler, and Understanding the dataset practitioners behind large Gert Breitfuss. 2024. How do professionals in smes language models. In Extended Abstracts of the CHI engage with ai and regulation? an interview study Conference on Human Factors in Computing Systems, in austria. In Proceedings of Mensch Und Computer CHI EA ’24, New York, NY, USA. Association for 2024, MuC ’24, page 646–650, New York, NY, USA. Computing Machinery. Association for Computing Machinery. Changle Qu, Sunhao Dai, Xiaochi Wei, Hengyi Cai, Jingfeng Yang, Hongye Jin, Ruixiang Tang, Xiao- Shuaiqiang Wang, Dawei Yin, Jun Xu, and Ji-Rong tian Han, Qizhang Feng, Haoming Jiang, Shaochen Wen. 2025. Tool learning with large language mod- Zhong, Bing Yin, and Xia Hu. 2024. Harnessing the els: A survey. Frontiers of Computer Science, power of llms in practice: A survey on chatgpt and 19(8):198343. beyond. ACM Trans. Knowl. Discov. Data, 18(6). Shezheng Song, Xiaopeng Li, Shasha Li, Shan Zhao, Zeyu Zhang, Xiaohe Bo, Chen Ma, Rui Li, Xu Chen, Jie Yu, Jun Ma, Xiaoguang Mao, Weimin Zhang, and Quanyu Dai, Jieming Zhu, Zhenhua Dong, and Ji- Meng Wang. 2025. How to bridge the gap between Rong Wen. 2024. A survey on the memory mecha- modalities: Survey on multimodal large language nism of large language model based agents. arXiv model. IEEE Transactions on Knowledge and Data preprint arXiv:2404.13501. Engineering, pages 1–20.