Text Classification with Machine Learning in Official Statistics: Experiences from the AIML4OS Project
Conference
Proposal Description
Recent advances in artificial intelligence (AI) and machine learning (ML) have created new opportunities to automate text-classification tasks that are fundamental to official statistics, such as coding textual entries according to standardized classification systems (e.g., NACE, ISCO, and COICOP).
Within the European AIML4OS project, Work Package 10 (WP10) brings together eight National Statistical Institutes (NSIs) and five observer NSIs to collaborate on the use case “From text to code - Experiences and potential of the use of AI/ML for classifying and coding.” This Invited Paper Session is dedicated to this use case and features contributions from participating NSIs, showcasing their experiences, methodologies, and perspectives while highlighting the diverse focus areas explored within the working group.
“Smart use of Large Language Models for production-ready statistical classification coding” (National Statistical Institute of Spain, INE): Large Language Models (LLMs) represent a major advancement in natural language processing. However, translating their capabilities into robust, cost-effective and low-latency automatic coding tools for statistical classifications remains challenging, particularly beyond the prototyping stage. This presentation presents strategies to address these limitations.
“Application of Large Language Models for NACE Activity Classification” (Statistics Poland): This paper presents preliminary results of research conducted within the AIML4OS project on the use of Large Language Models (LLMs) and other ML methods for the automatic classification of company activities according to the Polish Classification of Activities (PKD/NACE) using textual descriptions of business activities. The Retrieval-Augmented Generation (RAG) framework has been implemented to enhance classification quality by retrieving relevant NACE-related information as context for LLMs. The proposed methodology can be used to improve also the quality of other classification, such as Classification of Individual Consumption According to Purpose (COICOP) or International Standard Classification of Occupations (ISCO).
“Machine learning and LLM-based approaches for NACE classification at Statistics Norway”: this presentation describes the recent work at Statistics Norway on NACE classification. First, we discuss the problem of backcoding in which a company's activity description and the assigned NACE 2.1 code are used to backcode to the previous NACE 2.0 standard, and we report production-level performance in several machine learning approaches. Next, we describe the use of LLMs to directly classify from textual descriptors of companies to NACE 2.1 codes and present results from various approaches to prompting.
“The Hierarchy Question” (Statistics Austria): Insights from Hierarchical Text Classification for Standardized Codes”: this presentation examines hierarchical classification models that exploit the structure of classification systems, evaluating their potential to improve coding accuracy.
“Machine Learning Workflow enhancement for automatic coding” (National Statistical Institute of Luxembourg, STATEC): In statistical production, it is common to classify free text descriptions into official statistical classifications (GSBPM 5.2). This work reports on the development of standardized ML pipelines designed to support multiple classification use cases within STATEC. The proposed framework establishes a common approach to data handling, model training, validation, deployment, and documentation, facilitating governance and reproducibility while enabling reuse across use cases.