[논문]New Two-Level L1 Data Cache Bypassing Technique for High Performance GPUs

Kim, Gwang Bok; Kim, Cheol Hong

doi:10.3745/jips.01.0062

New Two-Level L1 Data Cache Bypassing Technique for High Performance GPUs 원문보기

Journal of information processing systems, v.17 no.1, 2021년, pp.51 - 62

Kim, Gwang Bok (R&D Center 2, SFA Engineering) , Kim, Cheol Hong (School of Computer Science and Engineering, Soongsil University)

Abstract ▼ AI-Helper

On-chip caches of graphics processing units (GPUs) have contributed to improved GPU performance by reducing long memory access latency. However, cache efficiency remains low despite the facts that recent GPUs have considerably mitigated the bottleneck problem of L1 data cache. Although the cache miss rate is a reasonable metric for cache efficiency, it is not necessarily proportional to GPU performance. In this study, we introduce a second key determinant to overcome the problem of predicting the performance gains from L1 data cache based on the assumption that miss rate only is not accurate. The proposed technique estimates the benefits of the cache by measuring the balance between cache efficiency and throughput. The throughput of the cache is predicted based on the warp occupancy information in the warp pool. Then, the warp occupancy is used for a second bypass phase when workloads show an ambiguous miss rate. In our proposed architecture, the L1 data cache is turned off for a long period when the warp occupancy is not high. Our two-level bypassing technique can be applied to recent GPU models and improves the performance by 6% on average compared to the architecture without bypassing. Moreover, it outperforms the conventional bottleneck-based bypassing techniques.

주제어

표/그림 (6)

그림 Fig. 1. Miss rates of L1 data cache and warp occupancy.
그림 Fig. 2. Hardware modification for the proposed technique.
표 Table 1. System configuration
그림 Fig. 3. Performance of miss rate-based bypassing with different thresholds.
그림 Fig. 4. Performance comparison of bypassing techniques.
그림 Fig. 5. Access and Misses reduced by the proposed technique.

참고문헌 (17)

W. Jia, K. A. Shaw, and M. Martonosi, "MRPB: memory request prioritization for massively parallel processors," in Proceedings of 2014 IEEE 20th International Symposium on High Performance Computer Architecture (HPCA), Orlando, FL, 2014, pp. 272-283.
NVIDIA Corporation, "NVIDIA Tesla P100: GP100 Pascal Architecture," 2016 [Online]. Available: https://images.nvidia.com/content/pdf/tesla/whitepaper/pascal-architecture-whitepaper.pdf.
C. T. Do, J. M. Kim, and C. H. Kim, "Application characteristics-aware sporadic cache bypassing for high performance GPGPUs," Journal of Parallel and Distributed Computing, vol. 122, pp. 238-250, 2018.

상세보기
J. Zhang, Y. He, F. Shen, and H. Tan, "Memory-aware TLP throttling and cache bypassing for GPUs," Cluster Computing, vol. 22, no. 1, pp. 871-883, 2019.

상세보기
NVIDIA Corporation, "NVIDA Tesla V100 GPU architecture," 2017 [Online]. Available: http://images.nvidia.com/content/volta-architecture/pdf/volta-architecture-whitepaper.pdf.
M. Gebhart, S. W. Keckler, B. Khailany, R. Krashinsky, and W. J. Dally, "Unifying primary cache, scratch, and register file memories in a throughput processor," in Proceedings of 2012 45th Annual IEEE/ACM International Symposium on Microarchitecture, Vancouver, Canada, 2012, pp. 96-106.
X. Xie, Y. Liang, Y. Wang, G. Sun, and T. Wang, "Coordinated static and dynamic cache bypassing for GPUs," in Proceedings of 2015 IEEE 21st International Symposium on High Performance Computer Architecture (HPCA), Burlingame, CA, 2015, pp. 76-88.
C. T. Do, J. M. Kim, and C. H. Kim, "Early miss prediction based periodic cache bypassing for high performance GPUs," Microprocessors and Microsystems, vol. 55, pp. 44-54, 2017.

상세보기
X. Chen, L. W. Chang, C. I. Rodrigues, J. Lv, Z. Wang, and W. M. Hwu, "Adaptive cache management for energy-efficient GPU computing," in Proceedings of 2014 47th Annual IEEE/ACM International Symposium on Microarchitecture, Cambridge, UK, 2014, pp. 343-355.
A. Sethia, D. A. Jamshidi, and S. Mahlke, "Mascar: speeding up GPU warps by reducing memory pitstops," in Proceedings of 2015 IEEE 21st International Symposium on High Performance Computer Architecture (HPCA), Burlingame, CA, 2015, pp. 174-185.
G. Koo, Y. Oh, W. W. Ro, and M. Annavaram, "Access pattern-aware cache management for improving data utilization in GPU," in Proceedings of the 44th Annual International Symposium on Computer Architecture, Toronto, Canada, 2017, pp. 307-319.
J. Fang, X. Zhang, S. Liu, and Z. Chang, "Miss-aware LLC buffer management strategy based on heterogeneous multi-core," The Journal of Supercomputing, vol. 75, no. 8, pp. 4519-4528, 2019.

상세보기
M. Khairy, A. Jain, T. M. Aamodt, and T. G. Rogers, "A detailed model for contemporary GPU memory systems," in Proceedings of 2019 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS), Madison, WI, 2019, pp. 141-142.
NVIDIA Corporation, "NVIDIA GeForce GTX 1080," 2016 [Online]. Available: https://international.download.nvidia.com/geforce-com/international/pdfs/GeForce_GTX_1080_Whitepaper_FINAL.pdf.
Z. Jia, M. Maggioni, B. Staiger, and D. P. Scarpazza, "Dissecting the NVIDIA Volta GPU architecture via microbenchmarking," 2018 [Online]. Available: https://arxiv.org/abs/1804.06826.
M. Bari, L. Stoltzfus, P. Lin, C. Liao, M. Emani, and B. Chapman, "Is data placement optimization still relevant on newer GPUs?," 2018 [Online]. Available: https://www.osti.gov/servlets/purl/1489476.
A. Karki, C. P. Keshava, S. M. Shivakumar, J. Skow, G. M. Hegde, and H. Jeon, "Tango: a deep neural network benchmark suite for various accelerators," in Proceedings of 2019 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS), Madison, WI, 2019, pp. 137-138.

내보내기 구분	파일저장 인쇄 메일전송
구성항목	기본정보 상세정보 관리번호, 논문명, 저널/프로시딩명, 저자 , 발행년, 권, 호, 시작페이지, 끝페이지, 발행기관 관리번호, 논문명, 대등논문명, 저자 , 저널/프로시딩명, 발행기관, 발행년, 발행언어, 권, 호, 시작페이지, 끝페이지, ISBN, ISSN, 주제분야, 키워드, 초록(한글), 초록(영문), 저자(소속기관)
저장형식	Text(ASCII format) Excel format RefWorks Direct Export RIS format (for Reference Manager, ProCite, EndNote), Scholar's Aids, Mendeley
메일정보	받는사람 (필수) @ 보내는사람 (선택) @ 제목 내용 KISTI 검색결과 이메일 서비스
안내	총 건의 자료가 검색되었습니다. 다운받으실 자료의 인덱스를 입력하세요. (1-10,000) 검색결과의 순서대로 최대 10,000건 까지 다운로드가 가능합니다. 데이타가 많을 경우 속도가 느려질 수 있습니다.(최대 2~3분 소요) 다운로드 파일은 UTF-8 형태로 저장됩니다. 파일의 내용이 제대로 보이지 않을실 때는 웹브라우저 상단의 보기 -> 인코딩 -> 자동선택 여부를 확인하십시오. ~ Text(ASCII format) Excel format

연합인증

New Two-Level L1 Data Cache Bypassing Technique for High Performance GPUs 원문보기

Abstract ▼ AI-Helper

주제어

표/그림 (6)

표/그림 (6)

참고문헌 (17)

이 논문을 인용한 문헌

관련 콘텐츠

원문 보기

원문 URL 링크

오픈액세스(OA) 유형

연관된 기능

이 논문과 함께 이용한 콘텐츠

AI-Helper ※ AI-Helper는 오픈소스 모델을 사용합니다.

선택된 텍스트

연합인증

New Two-Level L1 Data Cache Bypassing Technique for High Performance GPUs 원문보기

Abstract ▼ AI-Helper

주제어

표/그림 (6) 모든 표/그림 보기

표/그림 (6) 슬라이드로 보기

참고문헌 (17)

이 논문을 인용한 문헌

관련 콘텐츠

원문 보기

원문 URL 링크

오픈액세스(OA) 유형

연관된 기능

이 논문과 함께 이용한 콘텐츠

AI-Helper ※ AI-Helper는 오픈소스 모델을 사용합니다.

선택된 텍스트

표/그림 (6)

표/그림 (6)