| Nome: | Descrição: | Tamanho: | Formato: | |
|---|---|---|---|---|
| Preview-Abstrac | 154.95 KB | Adobe PDF |
Orientador(es)
Resumo(s)
This paper reports the adaptation of the Multidimensional Multiscale Parser (MMP) algorithm to CUDA. Specifically, we focus on memory optimization issues, such as the layout of data structures in memory, the type of GPU memory - shared, constant and global - and on achieving coalesced accesses. MMP is a demanding lossy compression algorithm for images. For example, MMP requires nearly 9000 seconds to encode the 512 x 512 Lenna image on a 2013's Intel Xeon. One of the main challenges to adapt MMP to manycore is related to the dependency over a pattern codebook which is built during the execution. This forces the input image to be processed sequentially. Nonetheless, CUDA-MMP achieves a 12x speedup over the sequential version when ran on an NVIDIA GTX 680. By further optimizing memory operations, the speedup is pushed to 17.1x.
Descrição
Computational Science and Its Applications - ICCSA 2014
14th International Conference, Guimarães, Portugal, June 30 - July 3, 204, Proceedings, Part IV.
Palavras-chave
CUDA image compression manycore computing memory optimization
Contexto Educativo
Citação
Domingues, P., Silva, J., Ribeiro, T., Rodrigues, N. M. M., De Carvalho, M. B., & De Faria, S. M. M. (2014). Optimizing Memory Usage and Accesses on CUDA-Based Recurrent Pattern Matching Image Compression. In Computational Science and Its Applications – ICCSA 2014, LNCS vol. 8582, 560-575. Springer, Cham. https://doi.org/10.1007/978-3-319-09147-1_41
Editora
Springer Nature
Coleções
Licença CC
Sem licença CC
