| Xie Xiaoyan(谢晓燕),Luo Xing,Zhang Heng,Zhou Enyang.[J].高技术通讯(英文),2026,32(3):240~249 |
|
| APR-BP: an adaptive memory partitioning and data reuse method for efficient back propagation |
| |
| DOI:10. 3772 / j. issn. 1006-6748. 2026. 03. 003 |
| 中文关键词: |
| 英文关键词: back-propagation, memory access scheduling, dynamic memory partitioning, data reuse |
| 基金项目: |
| Author Name | Affiliation | | Xie Xiaoyan(谢晓燕) | (School of Computer, Xi’an University of Posts and Telecommunications, Xi’an 710121, P. R. China) | | Luo Xing | | | Zhang Heng | | | Zhou Enyang | |
|
| Hits: 64 |
| Download times: 67 |
| 中文摘要: |
| |
| 英文摘要: |
| The access latency of off-chip memory is substantially higher than that of on-chip computation ,
especially during back propagation ( BP) . The limited on-chip buffers are strained by the frequent
reuse and write-back of activations , weights , and gradients , which necessitates repeated off-chip da-
ta transmission , and creates a severe performance bottleneck during training. To mitigate this , an
adaptive memory partitioning and data reuse method for BP(APR-BP) is proposed. Considering the
significant variation in access frequency and cross-layer reuse of activations , a priority-aware dynam-
ic memory partitioning ( PADP) strategy is introduced to characterize data reuse importance. By
PADP , activations with higher reuse potential are preferentially allocated to contiguous on-chip mem-
ory space , and effectively reduces redundant data movement and improves data reuse efficiency. Ai-
ming at single fixed reuse pattern cannot simultaneously balance on-chip utilization and data access
efficiency , a unified model is introduced to evaluate the benefit of three reuse patterns-input , weight
and output. The optimal pattern for each stage and layer is selected through benefit-aware optimiza-
tion. Experimental results demonstrate significant performance gains across three compute platforms ,
field-programmable gate array(FPGA) , graphics processing unit(GPU) , and embedded system-on-
chip(SoC) . Experimental results show that , compared to the baseline methods , the average data re-
use rate is improved by approximately 43. 0% , the total double data rate synchronous dynamic ran-
dom access memory(DDR) transfer volume is reduced by about 41. 5% , and on-chip buffer utiliza-
tion is increased by an average of 38. 5% with the proposed APR-BP. Furthermore , the execution
latency of BP for various networks is accelerated by an average of 1. 47 × . |
|
View Full Text
View/Add Comment Download reader |
| Close |
|
|
|