To cite this paper use one of the standards below:
The digitalization of public services has generated complex administrative documents, making Question Answering (QA) interfaces essential for citizen engagement. Retrieval-Augmented Generation (RAG) has emerged as the standard framework to ground these systems in proprietary knowledge bases while mitigating hallucinations. This paper evaluates three open-source embedding models, namely qwen3-embedding, bge-m3-embedding, and embedding-gemma, against proprietary OpenAI baselines within the Brazilian government's digital services catalog. We implemented four retrieval architectures: Dense, Hybrid (BM25 with Reciprocal Rank Fusion), and re-ranked variants using a cross-encoder. Evaluation used a synthetic dataset of 1,200 questions generated via the RAGAS framework, covering single-hop and multi-hop reasoning scenarios. Benchmarks show that open-source models with hybrid re-ranking achieve Recall@3 ≈ 0.82 and Recall@10 ≈ 0.93, matching proprietary performance. These findings demonstrate that open-source architectures are a technically viable and sovereign alternative for large-scale digital government initiatives.
With nearly 200,000 papers published, Galoá empowers scholars to share and discover cutting-edge research through our streamlined and accessible academic publishing platform.
Learn more about our products:
This proceedings is identified by a DOI , for use in citations or bibliographic references. Attention: this is not a DOI for the paper and as such cannot be used in Lattes to identify a particular work.
Check the link "How to cite" in the paper's page, to see how to properly cite the paper