Sovereign Information Retrieval: Benchmarking Open-Source Embedding Models for Brazilian Digital Public Services

Vol 57, 2025 - 341052
Complete Articles (CA)
Favorite this paper
How to cite this paper?
Abstract

The digitalization of public services has generated complex administrative documents, making Question Answering (QA) interfaces essential for citizen engagement. Retrieval-Augmented Generation (RAG) has emerged as the standard framework to ground these systems in proprietary knowledge bases while mitigating hallucinations. This paper evaluates three open-source embedding models, namely qwen3-embedding, bge-m3-embedding, and embedding-gemma, against proprietary OpenAI baselines within the Brazilian government's digital services catalog. We implemented four retrieval architectures: Dense, Hybrid (BM25 with Reciprocal Rank Fusion), and re-ranked variants using a cross-encoder. Evaluation used a synthetic dataset of 1,200 questions generated via the RAGAS framework, covering single-hop and multi-hop reasoning scenarios. Benchmarks show that open-source models with hybrid re-ranking achieve Recall@3 ≈ 0.82 and Recall@10 ≈ 0.93, matching proprietary performance. These findings demonstrate that open-source architectures are a technically viable and sovereign alternative for large-scale digital government initiatives.

Share your ideas or questions with the authors!

Did you know that the greatest stimulus in scientific and cultural development is curiosity? Leave your questions or suggestions to the author!

Sign in to interact

Have a question or suggestion? Share your feedback with the authors!

Institutions
  • 1 Universidade Federal de Alagoas
  • 2 Universidade Federal de Santa Catarina
  • 3 Universidade do Vale do Itajaí
Track
  • IA- OR and AI
Keywords
Natural Language Processing
Retrieval-Augmented Generation (RAG)
Information Retrieval
Public Services
Digital Government