To cite this paper use one of the standards below:
Deploying Large-Scale Language Models (LLMs) at the edge poses reliability challenges. This paper investigates the performance and fault tolerance in local inference of the Llama-3.2-1B-Q8 model on a Raspberry Pi 4B. Through experiments varying concurrency and context size (SQuAD v2.0), we analyzed resource saturation and latency. The results show that while RAM remains stable via 8-bit quantization, the CPU undergoes early saturation. This exhaustion causes performance failures, with high latency and instability under load. We reveal a paradox of overload in short requests (128 tokens) and the collapse of predictability in long contexts (1024 tokens). The study provides empirical evidence on the behavior of LLMs in restricted environments, supporting scaling mechanisms.
With nearly 200,000 papers published, Galoá empowers scholars to share and discover cutting-edge research through our streamlined and accessible academic publishing platform.
Learn more about our products:
This proceedings is identified by a DOI , for use in citations or bibliographic references. Attention: this is not a DOI for the paper and as such cannot be used in Lattes to identify a particular work.
Check the link "How to cite" in the paper's page, to see how to properly cite the paper