66th ISI World Statistics Congress

66th ISI World Statistics Congress

Investigating Re-Ranking Bottlenecks in Large-Scale Retrieval-Augmented Generation Systems and Proposing Performance Enhancement Techniques

Category: International Statistical Institute

Proposal Description

Retrieval-Augmented Generation (RAG) is an advanced artificial intelligence framework that combines information retrieval techniques with Large Language Models (LLMs) to generate accurate, context-aware, and factually grounded responses. Unlike conventional language models that rely solely on knowledge acquired during training, RAG retrieves relevant information from external data sources and incorporates it into the response generation process. This capability has made RAG a popular solution for enterprise search, policy document analysis, historical archives, healthcare knowledge systems, and other knowledge-intensive applications. However, as the size of data repositories increases and retrieval operations become distributed across multiple indexes, scalability challenges emerge, particularly during the re-ranking stage of the RAG pipeline.
The re-ranking component is responsible for evaluating retrieved documents and selecting the most relevant information for final response generation. In large-scale distributed RAG environments, numerous retrieval nodes independently return candidate documents to a central coordinator. The coordinator then performs semantic re-ranking, often using computationally expensive transformer-based models. As the number of retrieved candidates grows, re-ranking becomes a major bottleneck, increasing latency, network traffic, computational cost, and overall system complexity.
This research proposes a comprehensive evaluation of re-ranking bottlenecks in large RAG systems and investigates methods to improve scalability and performance. The methodology involves analyzing existing literature on distributed RAG architectures, federated search, key-value cache optimization, vector database scalability, hardware acceleration, and quality-aware retrieval pipelines. Building on these studies, the research proposes three improvement approaches: multi-stage re-ranking using lightweight sub-coordinators, vector compression to reduce communication overhead, and distributed key-value caching to support local computations and reduce inference delays. Experimental evaluation will be conducted using the MS MARCO benchmark dataset. Performance metrics will include retrieval accuracy, faithfulness, groundedness, throughput, network utilization, and coordinator re-ranking latency.
The proposed framework offers several advantages. First, it can reduce coordinator workload by distributing re-ranking responsibilities across multiple layers. Second, compression techniques can lower network bandwidth requirements and improve system responsiveness. Third, distributed caching can minimize repeated computations and accelerate retrieval operations. Collectively, these enhancements can improve scalability, reduce response time, and support efficient deployment of RAG systems in big-data environments.
Despite these benefits, the proposed approaches also have limitations. Multi-stage re-ranking may introduce architectural complexity and additional synchronization overhead. Compression techniques may result in information loss that affects ranking quality. Distributed caching requires additional storage resources and cache management mechanisms. Furthermore, maintaining consistency across distributed components may become challenging as system size increases.
In conclusion, this study addresses a critical yet underexplored challenge in large-scale RAG systems by focusing on re-ranking bottlenecks. Through systematic evaluation and experimentation, the research seeks to identify practical techniques that balance retrieval quality with computational efficiency. The outcomes are expected to provide valuable insights for researchers and practitioners developing next-generation RAG architectures capable of supporting large, distributed, and data-intensive applications while maintaining high levels of accuracy and responsiveness.