To cite this paper use one of the standards below:
Computational finite difference methods for partial differential equations presents data stencil patterns, when calculating time steps in a discrete domain. Due memory access strategies, parallel implementation of these methods may benefit from data locality if the data is arranged accordingly. In this paper we propose a new data structure and memory access pattern that may reduce non-coalescent
memory access and code divergence when loading neighborhood data in a GPU architecture. In our strategy, we use subdomains and extra buffers in order to optimize global memory access by warps. Our method achieves up to 1.3 times faster results than a classical strategy for finite difference GPU implementation
With nearly 200,000 papers published, Galoá empowers scholars to share and discover cutting-edge research through our streamlined and accessible academic publishing platform.
Learn more about our products:
This proceedings is identified by a DOI , for use in citations or bibliographic references. Attention: this is not a DOI for the paper and as such cannot be used in Lattes to identify a particular work.
Check the link "How to cite" in the paper's page, to see how to properly cite the paper