Current Issue

  • Select all
    |
    Cover Articl
  • Cover Articl
    Yang Yuze, Huang Zhengwei, Wang Yongwen
    Download PDF ( ) HTML ( )   Knowledge map   Save

    Irregular data access patterns in high-performance computing and intelligent computing often render traditional data prefetching techniques ineffective. Existing models that rely on fixed rules or offline learning based on specific program contexts also struggle to adapt to dynamically changing memory access patterns during runtime. While the Pythia reinforcement learning (RL) prefetching framework demonstrates adaptability through online learning, it still requires manual tuning under extreme irregular workloads, limiting its generalization in practical applications. This paper proposes IEP (Irregular Enhanced Pythia), a context-aware reinforcement learning prefetching framework to enhance the prediction capability for irregular memory access patterns. The framework introduces two key innovations. The first is an irregular feature enhancement module. It incorporates address bit masks and access sequence distance as state features to capture hidden spatiotemporal patterns in memory allocator behavior, thereby improving the representation of irregular memory accesses. The second is a hierarchical reward strategy module, which employs a dynamic reward mechanism combining confidence awareness and bandwidth sensitivity. This mechanism finely guides the agent’s learning process, accelerating policy optimization and improving final performance. Experiments were conducted using the ChampSim simulator, to test various irregular workloads. Results show that compared to the Pythia framework, the proposed solution achieves a maximum improvement of 2.27% in average prefetching accuracy and 2.90% in average single-core IPC for typical irregular workloads such as Ligra and PARSEC, while maintaining stable performance in multi-core environments.

  • Paper
  • Paper
    Wu Liangshun, Zhang Bin
    Download PDF ( ) HTML ( )   Knowledge map   Save

    We design an application-specific instruction-set processor (ASIP) acceleration core based on RISC-V for spiking neural network (SNN) computation. Unlike highly specialized FPGA or memristor/PIM solutions, our approach extends the RISC-V ISA to fuse SNN hotspots—synaptic vector dot-products and LIF/IZH neuron-state updates—at the instruction level, while reusing the GCC toolchain and scaling on multicore RISC-V platforms. The core integrates SIMD operations, fixed-point modules, and an event-driven multiplier, and introduces dedicated instructions (e.g., p.lif and p.izh). Evaluation on the cycle-accurate GVSoC virtual platform shows up to 14.6 times single-core speedup on MNIST-MLP and near-linear scaling from 1 to 16 cores, with accuracy close to the software baseline.

  • Paper
    Chen Zhihong, Du Yuan, Du Li
    Download PDF ( ) HTML ( )   Knowledge map   Save

    Ensuring the structural integrity of bridge expansion joints is critical for maintaining traffic safety and prolonging infrastructure lifespan. However, conventional inspection methods are generally labor-intensive, time-consuming, and disruptive to traffic flow. To overcome these limitations, this study proposes an edge-cloud collaborative acoustic monitoring system for the long-term health assessment of bridge expansion joints. Over an 18-month monitoring period spanning multiple bridges, approximately 21 000 acoustic signature samples of bridge expansion joints were collected. To ensure the accuracy and reliability, all samples were meticulously annotated by experienced highway engineers. Based on this dataset, an adaptive edge-cloud collaborative classification framework was developed. Specifically, the edge device employs a cascade classification model based on Support Vector Machines (SVM) for real-time, lightweight inference, while the cloud leverages a deep learning model based on Gated Recurrent Units (GRU) to perform more complex analysis. The cascaded classification model deployed on the edge device achieved an accuracy of 96.7%, with a 95% confidence interval (CI) of (0.963 2, 0.966 8). In comparison, the GRU-based classifier running on the cloud attained a higher accuracy of 98.2%, with a 95% CI of (0.978 3, 0.982 3). Furthermore, the proposed adaptive two-stage classification strategy reduced data transmission to less than 5% of the total collected data. These experimental results demonstrate that the proposed system offers a reliable, efficient, and accurate solution for the acoustic health monitoring of bridge expansion joints.

  • Paper
    Chen Cheng, Chen Guangwei, Chen Wentao, Sun Chenyang
    Download PDF ( ) HTML ( )   Knowledge map   Save

    Aiming at the problems of redundant volume, large link loss, and weak anti-electromagnetic interference of traditional FMCW radar transceiver systems in miniaturized devices, this paper designs and implements a highly integrated FMCW radar transceiver SiP module operating in the 4.2~4.4 GHz frequency band, which is suitable for aircraft altimeters. Based on System-in-Package (SiP) technology, the module adopts a 4-layer BT substrate to realize heterogeneous integration of multiple chips. It internally integrates a transmission link, a reception link, and a clock generator. A DDS is used to generate 500~700 MHz signals, which are then upconverted to the target frequency band in one step. The module features transmit power control, receive gain adjustment, and frequency-swept waveform configuration. The system is designed with a transmit output power of 20 dBm, a maximum receive link gain of 80.5 dB, and a transmit-receive isolation of ≥90 dB, with a compact package size of only 14 mm×14 mm. The test results show that the module has a transmit power flatness of less than 1 dB, a noise figure of only 4.08 dB at maximum gain, and an FMCW signal linearity of 0.012%. It can achieve kilometer-level altitude detection and fully meets the application requirements of aircraft for miniaturization, low power consumption, and high signal stability.

  • Paper
    Zou Can, Xun Taoyuan, Cai Jun, Zhong Jiajun, Zhu Yuehong, Zou Wanghui
    Download PDF ( ) HTML ( )   Knowledge map   Save

    This paper designs and implements an (Fast Fourier Transform,FFT) hardware accelerator and host computer system for bridge structural health monitoring. The hardware adopts a sequential 4-base FFT architecture with single-butterfly multiplexing to perform computations, effectively reducing resource overhead and power consumption while ensuring functional integrity. It supports configurable FFT sizes of 4, 16, 64 and 256 points. To enable data interaction and visualization, a host computer platform was further developed. This platform facilitates parameter configuration, operational control, and real-time display and analysis of frequency-domain results from the accelerator. The hardware accelerator has been verified on a CMOS 180 nm process and maintains stable operation at a 100 MHz operating frequency. Applied to bridge vibration signal processing, this system accurately extracts primary frequency components, meeting the comprehensive requirements of real-time performance, precision, and energy efficiency for bridge structural health monitoring.

  • Paper
    Wang Chao, Li Yongrui, Zhang Zhihan
    Download PDF ( ) HTML ( )   Knowledge map   Save

    Real-time imaging processing of spaceborne Synthetic Aperture Radar (SAR) requires multi-mode and large-scale computation under strictly limited on-board resources and harsh space radiation environments. Developing a dedicated SAR imaging processing System-on-Chip (SoC) chip, which achieves algorithm acceleration through specialized computing engine circuits, can effectively improve the efficiency of real-time imaging processing. During the design process of the dedicated chip, it is necessary to verify its functional boundaries and abnormal scenarios, and how to improve the verification coverage of specialized algorithm acceleration circuits is a core challenge. This paper proposes a verification method for SAR imaging processing chips based on fine-grained circuit matching modeling. An independent fine-grained reference model is established for each computing engine, and feature modeling such as circuit operation precision and circuit execution priority is added on the basis of realizing operational functions. This aims to solve the problem that traditional reference models cannot support data comparison for SAR computing engines. On this basis, a reusable and extensible verification environment is built based on the Universal Verification Methodology (UVM), adopting a dual-combination verification strategy of algorithm function verification and random configuration testing. Test results show that the proposed method can improve the verification coverage of computing engines by 14%~30%, meeting the requirements of SAR real-time imaging processing applications.

  • Paper
    Mao Ruipeng, Sun Lanhao, Huang Fei, Wang Shaohao
    Download PDF ( ) HTML ( )   Knowledge map   Save

    The two-transistor capacitorless (2T0C) gain-cell embedded dynamic random access memory (eDRAM) offers long data retention and high potential for three-dimensional (3D) integration, making it a compelling candidate for high-density embedded storage applications. However, write-data uniformity in large-scale 2T0C arrays is susceptible to various degradation mechanisms, thus driving the need for precise memory-channel modeling to ensure reliability. However, the storage-node voltage (VSN) suffers from stage-dependent ambiguity due to nonlinear capacitance and coupling effects, preventing it from being a unique state descriptor. To overcome this, we propose a unified Z-channel model centered on the stored charge (QSN) to accurately describe both write and hold operations. By adopting QSN instead of VSN as the fundamental state descriptor, the proposed framework resolves inherent representational ambiguity and facilitates the direct quantification of three dominant degradation mechanisms: write-history dependence, array parasitics, and retention leakage. To validate its generality, comprehensive Monte Carlo simulations were conducted across 2T0C arrays fabricated in multiple technology nodes. The results show that scaling down amorphous-oxide-semiconductor field effect transistors (AOSFETs) effectively suppresses the write-history-dependency, improves write uniformity in large 2T0C arrays, and achieves 7 500 seconds data retention.

  • Paper
    Xu Haina, Zhang Xun
    Download PDF ( ) HTML ( )   Knowledge map   Save

    In light of the urgent demand for system-level health management in the domestic application of the current VPX architecture within high-reliability fields such as aerospace and defense electronics, this paper proposes and designs a domestic health management module based on the collaborative architecture of Phytium processor, FPGA and MCU, and deeply integrates it into the VPX system framework to achieve comprehensive, intelligent, and full-life-cycle monitoring and management of the system operating status. Based on the VITA46.0 specification and IPMI protocol, the health management module establishes a three-level collaborative, clearly divided, and mutually redundant core control architecture of "Phytium processor+FPGA+MCU" through the collaboration of three core units, which significantly improves system reliability and fault tolerance, and reduces operation and maintenance costs. The Phytium D2000/8 undertakes system-level health status decision-making, data fusion processing, and global scheduling; the FPGA, as the acceleration control unit, is responsible for high-speed data acquisition, precise timing control, and custom interface expansion. The MCU focuses on real-time collection of key parameters such as voltage and temperature, local fault early warning, IPMI bus, and Ethernet communication. The module integrates domestic sensors, communication interfaces, and protocols, and realizes health management functions such as status monitoring, fault diagnosis, and fault prediction of key components in the VPX system through software-hardware collaboration. Through modular design, standardized interfaces, and custom protocols, the module can be seamlessly embedded into existing VPX systems, improving system operation and maintenance efficiency and localization rate. This solution fully adopts localized components and technical routes, has independent intellectual property rights, meets the requirements of the national independent and controllable strategy, and provides a replicable and promotable technical paradigm for the localization, intelligence, and sustainable operation and maintenance of high-performance embedded systems under the VPX architecture. The experimental and application results show that the health management module exhibits high reliability and accurate monitoring capability in harsh environments, and has significant application value.

  • Paper
    Tang Junlong, Dai Qiliang, Yin Zhenglin, Hou Gen
    Download PDF ( ) HTML ( )   Knowledge map   Save

    Aiming at the problems such as the limited on-chip resources of the self-developed DSP (Digital Signal Processor) and the insufficient storage space for large-scale data and programs, a configurable External Memory Interface (EMIF) design scheme is proposed and designed. This scheme achieves flexible access to three types of memory, namely asynchronous memory, SDRAM and SBSRAM, through four configurable chip selection space registers, and realizes efficient data transmission between the CPU and external memory by using an enhanced direct memory access module. EMIF functionality was verified using a testbench by comparing the consistency between the read and written data. The experimental results show that this design realizes the read-write burst access of 8/16/32-bit data. Under the 40 nm low-threshold process and an operating frequency of 200 MHz, the area is 25 119.360 2 μm2 and the power consumption is 1.296 mW.

  • Paper
    Liu Jixiang, Jiang Yingdan, Li Kun, Chen Wentao
    Download PDF ( ) HTML ( )   Knowledge map   Save

    Addressing the capacitor mismatch issue in high-precision successive approximation analog-to-digital converters (ADCs), this paper designs a foreground calibration technique based on sine signal input. By collecting kernel data for multiple fitting to ensure the signal-to-noise ratio (SNR) meets the specifications, the capacitor mismatch register values are obtained and OTP programming is performed. This effectively improves conversion accuracy and SNR without affecting the ADC sampling rate. This foreground calibration technique is derived from the Least Mean Squares (LMS) algorithm. It collects 16 KB kernel data from the SAR ADC and performs nonlinear least squares fitting using MATLAB. Drawing on the idea of the LMS algorithm, the residual signal undergoes multiple iterations, with each iteration adjusting the weight of each bit of the ADC accordingly. After approximately 1000 iterations, the SNR reaches 88 dB and the spurious-free dynamic range (SFDR) is 98 dB, which are 22 dB and 17 dB higher than before calibration, respectively. The simulation and test results show that this calibration technique effectively enhances the output performance of the ADC.

  • Paper
    Tang Hao
    Download PDF ( ) HTML ( )   Knowledge map   Save

    This paper presents an arbitrary interference waveform signal generator implemented based on an FPGA chip and the wideband DAC chip AD9739A. The generator fully utilizes the abundant logic, RAM, DSP, and high-speed interface resources of the FPGA to realize functional modules such as Gigabit/10-Gigabit Ethernet, an ARM soft-core processor, and high-speed cache, thereby achieving high-speed communication and data processing capabilities. To meet the requirements for rapid switching of wideband arbitrary modulated signal waveforms and to adapt to the high-speed DAC chip AD9739A with a sampling rate of up to 2.5 GSa/s, the design employs up to 20 low-speed parallel DDS IP cores inside the FPGA to generate waveform data. The data from these twenty cores is then sequentially output using a high-speed clock. This process requires meticulous design of the data flow and high-speed data read/write timing to realize the arbitrary interference waveform signal DAC functionality. This signal generator has been successfully applied in the project and has achieved excellent results.