CAWA: Coordinated warp scheduling and cache prioritization for critical warp acceleration of GPGPU workloads

Shin Ying Lee; Akhil Arunkumar; Carole-Jean Wu

doi:10.1145/2749469.2750418

CAWA: Coordinated warp scheduling and cache prioritization for critical warp acceleration of GPGPU workloads

Shin Ying Lee, Akhil Arunkumar, Carole-Jean Wu

Research output: Chapter in Book/Report/Conference proceeding › Conference contribution

65 Scopus citations

Abstract

The ubiquity of graphics processing unit (GPU) architectures has made them efficient alternatives to chip-multiprocessors for parallel workloads. GPUs achieve superior performance by making use of massive multi-threading and fast context-switching to hide pipeline stalls and memory access latency. However, recent characterization results have shown that general purpose GPU (GPGPU) applications commonly encounter long stall latencies that cannot be easily hidden with the large number of concurrent threads/warps. This results in varying execution time disparity between different parallel warps, hurting the overall performance of GPUs - the warp criticality problem. To tackle the warp criticality problem, we propose a coordinated solution, criticality-aware warp acceleration (CAWA), that efficiently manages compute and memory resources to accelerate the critical warp execution. Specifically, we design (1) an instruction-based and stall-based criticality predictor to identify the critical warp in a thread-block, (2) a criticality-aware warp scheduler that preferentially allocates more time resources to the critical warp, and (3) a criticality-aware cache reuse predictor that assists critical warp acceleration by retaining latency-critical and useful cache blocks in the L1 data cache. CAWA targets to remove the significant execution time disparity in order to improve resource utilization for GPGPU workloads. Our evaluation results show that, under the proposed coordinated scheduler and cache prioritization management scheme, the performance of the GPGPU workloads can be improved by 23% while other state-of-the-art schedulers, GTO and 2-level schedulers, improve performance by 16% and -2% respectively.

Original language	English (US)
Title of host publication	ISCA 2015 - 42nd Annual International Symposium on Computer Architecture, Conference Proceedings
Publisher	Institute of Electrical and Electronics Engineers Inc.
Pages	515-527
Number of pages	13
ISBN (Electronic)	9781450334020
DOIs	https://doi.org/10.1145/2749469.2750418
State	Published - Jun 13 2015
Event	42nd Annual International Symposium on Computer Architecture, ISCA 2015 - Portland, United States Duration: Jun 13 2015 → Jun 17 2015

Publication series

Name	Proceedings - International Symposium on Computer Architecture
Volume	13-17-June-2015
ISSN (Print)	1063-6897

Other

Other	42nd Annual International Symposium on Computer Architecture, ISCA 2015
Country/Territory	United States
City	Portland
Period	6/13/15 → 6/17/15

ASJC Scopus subject areas

Hardware and Architecture

Access to Document

10.1145/2749469.2750418

Cite this

Lee, S. Y., Arunkumar, A., & Wu, C.-J. (2015). CAWA: Coordinated warp scheduling and cache prioritization for critical warp acceleration of GPGPU workloads. In ISCA 2015 - 42nd Annual International Symposium on Computer Architecture, Conference Proceedings (pp. 515-527). (Proceedings - International Symposium on Computer Architecture; Vol. 13-17-June-2015). Institute of Electrical and Electronics Engineers Inc.. https://doi.org/10.1145/2749469.2750418

CAWA: Coordinated warp scheduling and cache prioritization for critical warp acceleration of GPGPU workloads. / Lee, Shin Ying; Arunkumar, Akhil; Wu, Carole-Jean.
ISCA 2015 - 42nd Annual International Symposium on Computer Architecture, Conference Proceedings. Institute of Electrical and Electronics Engineers Inc., 2015. p. 515-527 (Proceedings - International Symposium on Computer Architecture; Vol. 13-17-June-2015).

Research output: Chapter in Book/Report/Conference proceeding › Conference contribution

Lee, SY, Arunkumar, A & Wu, C-J 2015, CAWA: Coordinated warp scheduling and cache prioritization for critical warp acceleration of GPGPU workloads. in ISCA 2015 - 42nd Annual International Symposium on Computer Architecture, Conference Proceedings. Proceedings - International Symposium on Computer Architecture, vol. 13-17-June-2015, Institute of Electrical and Electronics Engineers Inc., pp. 515-527, 42nd Annual International Symposium on Computer Architecture, ISCA 2015, Portland, United States, 6/13/15. https://doi.org/10.1145/2749469.2750418

Lee SY, Arunkumar A, Wu CJ. CAWA: Coordinated warp scheduling and cache prioritization for critical warp acceleration of GPGPU workloads. In ISCA 2015 - 42nd Annual International Symposium on Computer Architecture, Conference Proceedings. Institute of Electrical and Electronics Engineers Inc. 2015. p. 515-527. (Proceedings - International Symposium on Computer Architecture). doi: 10.1145/2749469.2750418

Lee, Shin Ying ; Arunkumar, Akhil ; Wu, Carole-Jean. / CAWA : Coordinated warp scheduling and cache prioritization for critical warp acceleration of GPGPU workloads. ISCA 2015 - 42nd Annual International Symposium on Computer Architecture, Conference Proceedings. Institute of Electrical and Electronics Engineers Inc., 2015. pp. 515-527 (Proceedings - International Symposium on Computer Architecture).

@inproceedings{fbdf927997eb41a2806f6fc8a49633d1,

title = "CAWA: Coordinated warp scheduling and cache prioritization for critical warp acceleration of GPGPU workloads",

abstract = "The ubiquity of graphics processing unit (GPU) architectures has made them efficient alternatives to chip-multiprocessors for parallel workloads. GPUs achieve superior performance by making use of massive multi-threading and fast context-switching to hide pipeline stalls and memory access latency. However, recent characterization results have shown that general purpose GPU (GPGPU) applications commonly encounter long stall latencies that cannot be easily hidden with the large number of concurrent threads/warps. This results in varying execution time disparity between different parallel warps, hurting the overall performance of GPUs - the warp criticality problem. To tackle the warp criticality problem, we propose a coordinated solution, criticality-aware warp acceleration (CAWA), that efficiently manages compute and memory resources to accelerate the critical warp execution. Specifically, we design (1) an instruction-based and stall-based criticality predictor to identify the critical warp in a thread-block, (2) a criticality-aware warp scheduler that preferentially allocates more time resources to the critical warp, and (3) a criticality-aware cache reuse predictor that assists critical warp acceleration by retaining latency-critical and useful cache blocks in the L1 data cache. CAWA targets to remove the significant execution time disparity in order to improve resource utilization for GPGPU workloads. Our evaluation results show that, under the proposed coordinated scheduler and cache prioritization management scheme, the performance of the GPGPU workloads can be improved by 23% while other state-of-the-art schedulers, GTO and 2-level schedulers, improve performance by 16% and -2% respectively.",

author = "Lee, {Shin Ying} and Akhil Arunkumar and Carole-Jean Wu",

year = "2015",

month = jun,

day = "13",

doi = "10.1145/2749469.2750418",

language = "English (US)",

series = "Proceedings - International Symposium on Computer Architecture",

publisher = "Institute of Electrical and Electronics Engineers Inc.",

pages = "515--527",

booktitle = "ISCA 2015 - 42nd Annual International Symposium on Computer Architecture, Conference Proceedings",

note = "42nd Annual International Symposium on Computer Architecture, ISCA 2015 ; Conference date: 13-06-2015 Through 17-06-2015",

}

TY - GEN

T1 - CAWA

T2 - 42nd Annual International Symposium on Computer Architecture, ISCA 2015

AU - Lee, Shin Ying

AU - Arunkumar, Akhil

AU - Wu, Carole-Jean

PY - 2015/6/13

Y1 - 2015/6/13

N2 - The ubiquity of graphics processing unit (GPU) architectures has made them efficient alternatives to chip-multiprocessors for parallel workloads. GPUs achieve superior performance by making use of massive multi-threading and fast context-switching to hide pipeline stalls and memory access latency. However, recent characterization results have shown that general purpose GPU (GPGPU) applications commonly encounter long stall latencies that cannot be easily hidden with the large number of concurrent threads/warps. This results in varying execution time disparity between different parallel warps, hurting the overall performance of GPUs - the warp criticality problem. To tackle the warp criticality problem, we propose a coordinated solution, criticality-aware warp acceleration (CAWA), that efficiently manages compute and memory resources to accelerate the critical warp execution. Specifically, we design (1) an instruction-based and stall-based criticality predictor to identify the critical warp in a thread-block, (2) a criticality-aware warp scheduler that preferentially allocates more time resources to the critical warp, and (3) a criticality-aware cache reuse predictor that assists critical warp acceleration by retaining latency-critical and useful cache blocks in the L1 data cache. CAWA targets to remove the significant execution time disparity in order to improve resource utilization for GPGPU workloads. Our evaluation results show that, under the proposed coordinated scheduler and cache prioritization management scheme, the performance of the GPGPU workloads can be improved by 23% while other state-of-the-art schedulers, GTO and 2-level schedulers, improve performance by 16% and -2% respectively.

AB - The ubiquity of graphics processing unit (GPU) architectures has made them efficient alternatives to chip-multiprocessors for parallel workloads. GPUs achieve superior performance by making use of massive multi-threading and fast context-switching to hide pipeline stalls and memory access latency. However, recent characterization results have shown that general purpose GPU (GPGPU) applications commonly encounter long stall latencies that cannot be easily hidden with the large number of concurrent threads/warps. This results in varying execution time disparity between different parallel warps, hurting the overall performance of GPUs - the warp criticality problem. To tackle the warp criticality problem, we propose a coordinated solution, criticality-aware warp acceleration (CAWA), that efficiently manages compute and memory resources to accelerate the critical warp execution. Specifically, we design (1) an instruction-based and stall-based criticality predictor to identify the critical warp in a thread-block, (2) a criticality-aware warp scheduler that preferentially allocates more time resources to the critical warp, and (3) a criticality-aware cache reuse predictor that assists critical warp acceleration by retaining latency-critical and useful cache blocks in the L1 data cache. CAWA targets to remove the significant execution time disparity in order to improve resource utilization for GPGPU workloads. Our evaluation results show that, under the proposed coordinated scheduler and cache prioritization management scheme, the performance of the GPGPU workloads can be improved by 23% while other state-of-the-art schedulers, GTO and 2-level schedulers, improve performance by 16% and -2% respectively.

UR - http://www.scopus.com/inward/record.url?scp=84960122845&partnerID=8YFLogxK

UR - http://www.scopus.com/inward/citedby.url?scp=84960122845&partnerID=8YFLogxK

U2 - 10.1145/2749469.2750418

DO - 10.1145/2749469.2750418

M3 - Conference contribution

AN - SCOPUS:84960122845

T3 - Proceedings - International Symposium on Computer Architecture

SP - 515

EP - 527

BT - ISCA 2015 - 42nd Annual International Symposium on Computer Architecture, Conference Proceedings

PB - Institute of Electrical and Electronics Engineers Inc.

Y2 - 13 June 2015 through 17 June 2015

ER -

CAWA: Coordinated warp scheduling and cache prioritization for critical warp acceleration of GPGPU workloads

Abstract

Publication series

Other

ASJC Scopus subject areas

Access to Document

Other files and links

Fingerprint

Cite this