| Deng Junyong (邓军勇)*,Li Zhe*,Xie Xiaoyan**,Zhu Yun*,Lu Songtao*,Hu Bin***.[J].高技术通讯(英文),2026,32(3):322~332 |
|
| RSM: dynamic self-reconfiguration mechanism for AI accelerators based on runtime status monitoring |
| |
| DOI:10. 3772 / j. issn. 1006-6748. 2026. 03. 011 |
| 中文关键词: |
| 英文关键词: AI accelerators, runtime status monitoring, self-reconfiguration mechanism, programmable state manager |
| 基金项目: |
| Author Name | Affiliation | | Deng Junyong (邓军勇)* | (*School of Electronic Engineering, Xi’an University of Posts and Telecommunications, Xi’an 710121, P. R. China)
(**School of Computer, Xi’an University of Posts and Telecommunications, Xi’an 710121, P. R. China)
(***Department of Computer Science and Technology, Kean University, 07083, USA) | | Li Zhe* | | | Xie Xiaoyan** | | | Zhu Yun* | | | Lu Songtao* | | | Hu Bin*** | |
|
| Hits: 61 |
| Download times: 63 |
| 中文摘要: |
| |
| 英文摘要: |
| Driven by the increasing complexity of deep learning models , traditional processors face limita-
tions in terms of energy efficiency and flexibility. Coarse-grained reconfigurable arrays ( CGRAs )
balance performance and energy efficiency through dynamic configuration , but they lack real-time
task state monitoring and precise resource scheduling. This results in low processing element ( PE )
utilization and significant performance loss. To address these issues , this paper proposes a self-
reconfiguration mechanism based on runtime status monitoring (RSM) . A programmable state man-
ager performs runtime evaluation and decision-making for the PE array , enabling dynamic task
scheduling and context switching within 6 clock cycles. Evaluations on simultaneous localization and
mapping ( SLAM ) and expectation-maximization (EM) routing algorithms demonstrate the effective-
ness of the proposed RSM mechanism. The proposed design achieves an EM routing time of 9. 30 μs
and a SLAM execution time of 0. 13 ms. For the SLAM algorithm , the execution time is reduced
from 119 052 cycles to 52 050 cycles , corresponding to a 56. 28% reduction and a 2. 29 × speed-
up. For the EM routing algorithm , the execution time is reduced from 7 521 cycles to 3 730 cycles ,
corresponding to a 50. 41% reduction and a 2. 02 × speedup. In addition , the reconfiguration time is
reduced by 47. 35% for SLAM and 49. 97% for EM routing. PE utilization increases from 20. 00% to
40.00% for SLAM , and instruction-flow PE utilization increases from 25. 00% to 50. 00% for EM
routing. |
|
View Full Text
View/Add Comment Download reader |
| Close |
|
|
|