ATLAS: B ENCHMARKING AND A DAPTING LLM S FOR G LOBAL T RADE VIA H ARMONIZED TARIFF C ODE C LASSIFICATION Pritish Yuvraj Siva Devarakonda∗ Flexify.AI Flexify.AI pritish@flexify.ai siva@flexify.ai A BSTRACT arXiv:2509.18400v1 [cs.AI] 22 Sep 2025 Accurate classification of products under the Harmonized Tariff Schedule (HTS) is a critical bottleneck in global trade, yet it has received little attention from the machine learning community. Misclassification can halt shipments entirely, with major postal operators suspending deliveries to the U.S. due to incomplete customs documentation. We introduce the first benchmark for HTS code classification, derived from the U.S. Customs Rulings Online Search System (CROSS). Evaluating leading LLMs, we find that our fine-tuned ATLAS model (LLaMA-3.3-70B) achieves 40% fully correct 10-digit classifications and 57.5% correct 6-digit classifi- cations—improvements of +15 points over GPT-5-Thinking and +27.5 points over Gemini-2.5-Pro-Thinking. Beyond accuracy, ATLAS is 5× cheaper than GPT-5-thinking and 8× cheaper than Gemini-2.5-Pro-Thinking, and can be self-hosted to guarantee data pri- vacy—an essential requirement in high stake industries like Automotives ,Indus- trials, Semiconductors etc. for trade and compliance workflows. While ATLAS sets a strong baseline, the benchmark remains highly challenging, with only 40% 10-digit accuracy. By releasing both dataset and model, we aim to position HTS classification as a new community benchmark task. We invite future work in retrieval, reasoning and alignment to advance progress on this high-impact global trade problem. 1 1 I NTRODUCTION Every product imported into the global market must be assigned a Harmonized Tariff Schedule (HTS) code. These codes, standardized by the World Customs Organization (WCO), are ten digits long. The first six digits are harmonized across all participating countries, while the last four digits are country-specific extensions. Correctly identifying the first six digits enables global interoper- ability, while the full ten-digit code is required for compliance with U.S. customs. The HTS is deeply hierarchical: 22 sections are divided into 99 chapters, which expand into thou- sands of subheadings. Chapters 1–97 correspond to stable product categories (such as chemicals, machinery, and textiles), whereas Chapters 98–99 capture temporary and special provisions that change frequently. This structure makes tariff classification a natural hierarchical machine learning task, where six-digit accuracy captures worldwide consistency, and ten-digit accuracy reflects the U.S.-specific extension. ∗ Website: https://tariffpro.flexify.ai/ 1 1. HTS CROSS Rulings Dataset: https://huggingface.co/datasets/flexifyai/cross_rulings_hts_dataset_ for_tariffs 2. Atlas LLM Model: https://huggingface.co/flexifyai/atlas-llama3.3-70b-hts-classification 3. Atlas LLM Model Demo: https://flexifyai-atlas-llama3-3-70b-hts-demo.hf.space/?logs=container& __theme=system&deep_link=FFJuTJsv_fM 1 Despite its centrality, classification remains a major bottleneck. Recent trade policy changes, for example, the modifications to the de minimis exemption, require that any imported good valued above $100 must be assigned a valid HTS code. The HTS itself spans more than 17,000 pages of PDF documents, making manual assignment infeasible at scale. The consequences are global: in 2025, several major postal operators suspended parcel delivery to the United States, citing the inability to assign correct HTS codes and complete customs documentation tim (2025), reu (2025), usa (2025). More than thirty countries were affected, highlighting the fragility of global trade flows when classification is not available at scale. Large Language Models (LLMs) offer a scalable alternative. Their capacity for semantic reasoning and structured classification makes them natural candidates for HTS code classification, where fine- grained distinctions must be captured. Moreover, since the first six digits are harmonized worldwide, advances in HTS classification can simultaneously benefit global trade systems, while the U.S.- specific digits directly address compliance in the American market. 1.1 C ONTRIBUTIONS Our key contributions are: • We release the first open-source benchmark for HTS classification Yuvraj & Devarakonda (2025a), constructed from the U.S. Customs Rulings Online Search System (CROSS), in- cluding training, validation, and test splits. • We benchmark leading proprietary and open-source models, including GPT-5- Thinking OpenAI (2025a), Gemini-2.5-Pro-Thinking DeepMind (2025), LLaMA-3.3-70B Grattafiori et al. (2024), DeepSeek-R1 (05/28) DeepSeek-AI et al. (2025), and GPT-OSS- 120B OpenAI (2025b). • We fine-tuned LLaMA-3.3-70B with supervised fine-tuning (SFT) to create the specialized model ATLAS, which we open source Yuvraj & Devarakonda (2025b). ATLAS achieves 40% accuracy at the 10-digit level and 57.5% at the 6-digit level, substantially outper- forming GPT-5-Thinking (25%) and Gemini-2.5-Pro-Thinking (13.5%). • Beyond accuracy, ATLAS is significantly more cost-efficient—up to 8× cheaper per in- ference—and supports privacy-preserving deployment through self-hosting, ensuring that sensitive trade data never leaves secure environments. Together, these contributions establish tariff code classification as a new frontier for LLM evaluation, situated at the heart of compliance for Global Commerce and Trade. 1.2 PAPER ROADMAP The remainder of this paper is organized as follows. Section 2 describes the construction of our CROSS-based dataset and its transformation into a machine-learning–ready format. Section 3 de- tails the fine-tuning procedure for ATLAS. Section 4 presents evaluation results across multiple pro- prietary and open-source LLMs, analyzing both accuracy and cost efficiency. Finally, we conclude with a summary of key findings and future research directions in Section 5. 2 DATASETS A central contribution of this work is the construction of the first large-scale dataset for Harmonized Tariff Schedule (HTS) classification, derived from the U.S. Customs Rulings Online Search System (CROSS) Customs & Protection (2025). CROSS contains legally binding decisions issued by U.S. Customs and Border Protection (CBP), in which importers or brokers sought clarification on the correct HTS code for specific products. These rulings are authoritative, high-value examples of tariff classification in practice, but are lengthy, unstructured, and scattered across thousands of HTML pages—making them inaccessible for machine learning research. 2 2.1 DATA C OLLECTION We developed an automated browser agent Project (2025); Google (2025); Pirogov (2025) to sys- tematically scrape CROSS. Each ruling was matched to a 10-digit HTS code obtained from the official HTS U.S. website Commission (2025). After filtering and cleaning, the final dataset spans 18,731 rulings covering 2,992 unique HTS codes across a broad range of product categories. Not every HTS code appears in CROSS, since only disputed or clarified codes are documented. However, the presence of a code in CROSS is itself informative: frequent rulings signal categories that are high-demand or ambiguous in practice, while absent codes suggest stable or rarely used classifications. 2.2 DATA T RANSFORMATION INTO LLM-T RAINABLE F ORMAT Raw CROSS rulings are official letters—legalistic, verbose, and inconsistent in structure. To make them suitable for supervised learning, we transformed each ruling into a structured prompt–response format using GPT-4o-mini OpenAI et al. (2024). This lightweight model was cost-effective and sufficient for information extraction. Prompt template. Each ruling was converted into the following instruction format: Given the following HTS ruling information: HTS Code: {hts_code} Ruling Number: {ruling_number} Title: {title} Date: {date} URL: {url} Summary: {summary} Content: {content} Please analyze this information and provide: a) Create a product description that the user was initially getting the HTS US code f b) Create a reasoning path justifying why the HTS US code is correct c) Return the HTS US code Format your response as follows: User: What is the HTS US Code for [product_description]? Model: HTS US Code -> [HTS US Code] Reasoning -> [detailed_reasoning_path] This design forces models to both predict the code and provide a reasoning path, aligning with recent work on chain-of-thought reasoning Wei et al. (2023). 2.3 DATASET S PLITS From the 18,731 processed rulings, we randomly sampled 200 examples for validation and 200 for final testing, holding them out strictly from training. The remaining 18,254 rulings form the training set. This ensures a clean separation between model development and final evaluation. We have uploaded the dataset to hugging-face Yuvraj & Devarakonda (2025a). 2.4 D ISCUSSION This dataset poses unique challenges: (1) rulings are lengthy and often hinge on subtle distinctions (e.g., partially fabricated vs. fully fabricated semiconductor wafers); (2) correctness has a hierarchi- cal structure (6-digit vs. 10-digit); and (3) errors carry real-world consequences for trade and com- 3 Table 1: Distribution of the CROSS dataset into training, validation, and test splits. Split Number of Rulings Training 18,254 Validation 200 Test 200 pliance. By releasing this benchmark, we aim to establish tariff classification as a novel, high-impact evaluation task for LLMs, complementing existing benchmarks in reasoning, code generation, and multilingual understanding. 3 M ODEL T RAINING While several open-source large language models could, in principle, be adapted for tariff clas- sification, we made a deliberate and principled choice to focus exclusively on LLaMA-3.3-70B Grattafiori et al. (2024). Two factors motivated this decision. First, practical budget constraints made it infeasible to fine-tune multiple frontier models at scale. Second, LLaMA-3.3-70B is a dense architecture, making it both simpler to fine-tune and easier to deploy in inference settings com- pared to Mixture-of-Experts (MoE) architectures such as DeepSeek-R1 or GPT-OSS-120B. From a community perspective, providing a dense and reproducible baseline lowers the entry barrier for downstream research: training and inference pipelines are easier to set up, memory usage is more predictable, and accuracy is less sensitive to expert routing heuristics. 3.1 S UPERVISED F INE -T UNING O BJECTIVE We adapted LLaMA-3.3-70B to the CROSS dataset using supervised fine-tuning (SFT) Brown et al. (2020); Ouyang et al. (2022). Each ruling was transformed into an input–output pair, where the input is a ruling-derived product description and the output is the correct HTS code along with a reasoning trace. This makes the task well aligned with the SFT paradigm, which minimizes the token-level negative log-likelihood of ground-truth outputs. Formally, for an input sequence x = (x1 , . . . , xn ) and target sequence y = (y1 , . . . , ym ), the model with parameters θ defines conditional probabilities pθ (yt | x, y