Sovereign Information Retrieval: Benchmarking Open-Source Embedding Models for Brazilian Digital Public Services

Vol 57, 2025 - 341052
Trabalho completo (Oral)
Favoritar este trabalho
Como citar esse trabalho?
Resumo

The digitalization of public services has generated complex administrative documents, making Question Answering (QA) interfaces essential for citizen engagement. Retrieval-Augmented Generation (RAG) has emerged as the standard framework to ground these systems in proprietary knowledge bases while mitigating hallucinations. This paper evaluates three open-source embedding models, namely qwen3-embedding, bge-m3-embedding, and embedding-gemma, against proprietary OpenAI baselines within the Brazilian government's digital services catalog. We implemented four retrieval architectures: Dense, Hybrid (BM25 with Reciprocal Rank Fusion), and re-ranked variants using a cross-encoder. Evaluation used a synthetic dataset of 1,200 questions generated via the RAGAS framework, covering single-hop and multi-hop reasoning scenarios. Benchmarks show that open-source models with hybrid re-ranking achieve Recall@3 ≈ 0.82 and Recall@10 ≈ 0.93, matching proprietary performance. These findings demonstrate that open-source architectures are a technically viable and sovereign alternative for large-scale digital government initiatives.

Compartilhe suas ideias ou dúvidas com os autores!

Sabia que o maior estímulo no desenvolvimento científico e cultural é a curiosidade? Deixe seus questionamentos ou sugestões para o autor!

Faça login para interagir

Tem uma dúvida ou sugestão? Compartilhe seu feedback com os autores!

Instituições
  • 1 Universidade Federal de Alagoas
  • 2 Universidade Federal de Santa Catarina
  • 3 Universidade do Vale do Itajaí
Eixo Temático
  • PO&IA – Pesquisa Operacional com Inteligência Artificial
Palavras-chave
Natural Language Processing
Retrieval-Augmented Generation (RAG)
Information Retrieval
Public Services
Digital Government