Processing unit design for high-performance computing processors

LIU Yu, ZHANG Jie, ZHOU Le

Integrated Circuits and Embedded Systems ›› 2025, Vol. 25 ›› Issue (9) : 57-62.

PDF(8376 KB)
PDF(8376 KB)
Integrated Circuits and Embedded Systems ›› 2025, Vol. 25 ›› Issue (9) : 57-62. DOI: 10.20193/j.ices2097-4191.2025.0039
Research Paper

Processing unit design for high-performance computing processors

Author information +
History +

Abstract

This paper proposes a design for a compute-in-memory processing unit (CIMPU) tailored for high-performance computing processors. The CIMPU integrates a multi-precision arithmetic operator and on-chip storage, enabling computations to be performed locally without accessing external buses. A hardware pipeline is further designed based on the CIMPU's architectural characteristics to optimize processing efficiency. The operation unit design scheme proposed in this paper has a good performance-power-consumption ratio advantage. We evaluated the computational performance of the design, with a performance to power ratio of 2.47 TOPS/W@INT8. It is significantly superior to other similar processor architectures and is suitable for large-scale deployment as a high computing power processor core.

Key words

high-performance computing processor / processing unit / scalability / pipeline / FPGA

Cite this article

Download Citations
LIU Yu , ZHANG Jie , ZHOU Le. Processing unit design for high-performance computing processors[J]. Integrated Circuits and Embedded Systems. 2025, 25(9): 57-62 https://doi.org/10.20193/j.ices2097-4191.2025.0039

References

[1]
PAN J, ZHOU G. A Survey of Research in Large Language Models for Electronic Design Automation[C]// ACM Transactions on Design Automation of Electronic Systems, 2025:34-55.
[2]
HE T, CHEN X, WANG G. Research on Open Source Processor and Analysis of Current Development Dilemma Based on RISC-V[C]// International Conference on Computer and Communication Systems (ICCCS), 2023:768-774.
[3]
XU D, ZHANG H, LIU X. Fast On-device LLM Inference with NPUs[C]// ASPLOS'2025,2025:445-462.
[4]
SOLAIMAN M, SOLAIMAN G M. A Novel Approach to Design and Verification of a Pipelined Microprocessor with Hazard Detection and Stall Insertion[C]// International Conference on I-SMAC, 2024:308-312.
[5]
LIN C Y, YEN H T, HUNG C L. Efficient Strategies of Compressing Three-Dimensional Sparse Arrays Based on Intel XEON and Intel XEON Phi Environments[C]// IEEE International Conference on Computer and Information Technology, 2015:1383-1388.
[6]
SATO M, TSUJI M. OpenACC Execution Models for Manycore Processor with ARM SVE[C]// Proceedings of the HPC Asia 2023 Workshops (HPCAsia '23 Workshops), 2023:73-77.
[7]
洪一, 方体莲. “魂芯一号”数字信号处理器及应用[J]. 中国科学, 2015, 45(4):574-586.
HONG Y, FANG T L. "Hunxin-1" Digital Signal Processor and its Applications[J]. Scientia Sinica, 2015, 45(4):574-586 (in Chinese).
[8]
LAKHDAR D M, ABDELKRIM M, MOHAMMED D. Optimisation of the implementations in fixed and floating point of tracking algorithm on DSP of C6000 of TI’s familly[C]// International Conference on Communications, Control Systems and Signal Processing, 2020:254-259.
[9]
OU S H, CHO Y, LIU C W. Improving datapath utilization of programmable DSP with composite functional units[C]// IEEE International Symposium on Circuits and Systems (ISCAS),Seattle,WA,USA, 2008:3438-3441.
[10]
CHOQUETTE J, GANDHI W, GIROUX O, et al. NVIDIA A100 Tensor Core GPU: Performance and Innovation[J]. IEEE Micro, 2021:29-35.
[11]
ZHU M, ZHANG T, GU Z, et al. Sparse Tensor Core: algorithm and hardware co-design for vector-wise sparse neural networks on modern gpus[C]// 52nd Annual IEEE/ACM International Symposium on Microarchitecture (MICRO), 2019:127-142.
[12]
LEE S, KIM K, KWAK J, et al. An Efficient NPU-Aware Filter Pruning in Convolutional Neural Network[C]// International Conference on Electronics,Information,and Communication (ICEIC), 2023:1-3.
[13]
JAVAID A, AHMED T, ALI S. Performance Evaluation of Xilinx Zynq UltraScale+ RFSoC Device for Low Latency Applications[C]// International Bhurban Conference on Applied Sciences and Technology, 2022:1041-1046.
[14]
XILINX. Ultra-Low Latency and High-Speed RF Data Converter Technology[EB/OL].[2025-06]. https://edit.wpgdadawant.com/uploads/news_file/blog/2021/3759/tech_files/sv-1621827910.pdf.
PDF(8376 KB)

Accesses

Citation

Detail

Sections
Recommended

/