Home Browse Just Accepted

Just Accepted

Note: The articles listed below have been peer-reviewed and accepted for publication in this journal. These articles have not yet been scheduled for a specific issue; their content and layout may undergo minor changes in the final published version. Please refer to the final published version as the definitive one. This journal has assigned each of these articles a unique and persistent DOI. You may use the DOI to cite this article directly.
Please wait a minute...
  • Select all
    |
  • Feng Xuning, Zhang Deming, Hu Yuanqi, Zhu Dapeng, Cheng Yuanqing
    Accepted: 2026-08-03
    Angular encoder chips often operate in complex environments and impose strict requirements on signal filtering. However, constrained by limited on-chip space, integrating complex filtering circuits remains a significant challenge. To address this issue, this article proposes a folded Kalman filter circuit that can co-optimize area and latency. The design employs four optimization strategies—a folded architecture, filter model splitting, Kalman parameter pruning, and operation sequence reordering—which significantly reduce hardware overhead. Processing a single-channel signal requires only 2 adders, 1 multiplier, and 1 divider. Implemented in 180nm process node at a 50MHz clock frequency, the circuit supports high-speed signal processing with a sampling rate of up to 2MSPS, achieving high-frequency noise attenuation up to 22dB. Without compromising filtering accuracy, our design reduces DSP resource consumption by 68% and computational latency by 53% compared to related works. It can be effectively applied to miniature signal processing chips, such as angular encoders.
  • Accepted: 2026-08-03
    【目的】目前如光传输网络OTN(Optical Transmission Network)、以太网数据通信传输的研究热点都集中在高速率传输,例如40Gbps以上,相应的串行器/解串器SerDes(Serializer-Deserializer)技术和芯片也随之推出,这些高速SerDes都存在下限速率,目前SerDes芯片存在速率下限,通常在500Mbps左右,因此500Mbps以下的数据串行传输不能直接使用SerDes,为了解决此问题,本文深入分析了过采样与SerDes结合的数据传输,并重点研究了接收端的过采样数据恢复。【方法】本文给出了基于SerDes的时域过采样设计,并详细设计了接收端的过采样数据恢复模块,给出了设计接口和信号说明;针对发送端、接收端的时钟频差时的过采样给出了详细分析和采样补偿或者舍弃,而且讨论了长连1或者长连0的最大容许长度,以便选择何种物理层编码码型。最后,对两种频差情况进行了详细仿真和讨论,并且基于紫光国产FPGA芯片PG2T100的板卡开展了测试验证。【结果】仿真和测试结果表明,基于SerDes的时域过采样设计系统满足500Mbps以下的数据传输需求。【结论】基于SerDes的时域过采样系统设计详细、仿真全面、上板测试充分,为500Mbps以下的数据串行传输提供了工程设计指导。
  • Accepted: 2026-07-31
    面向低空起降点、无人机物流节点和园区周界等夜间安防场景,针对单目旋转摄像头周期扫描引起的目标短时不可见、低照度条件下单一视觉线索退化及跨扇区身份连续维护困难,本文构建了一种夜间全域扫描多模态检测跟踪系统。系统采用四扇区时序覆盖,将颜色、纹理和形状互补表征、自适应加权融合、质量门控、模板更新、跨扇区ID继承及可视化回溯纳入统一闭环。在90°步进角和113°水平视场角条件下,相邻扇区形成约23°重叠区,并在15 s内完成一次360°周期覆盖。融合模块依据归一化置信度分配模态权重,质量门控进一步抑制低置信度模态贡献,并控制每帧模板更新。实验结果表明,系统CLE为28.64、OS为0.574、Precision为0.800、Success为0.720,端到端FPS为10.45。移除纹理模态、PSR/APCE和质量门控后,Success分别降至0.599、0.612和0.652。20个有效跨扇区目标中,18个正确继承原ID,ID保持率和恢复成功率均为90%,共出现2次ID切换和2次误匹配,验证了系统在夜间大视场监控和低空基础设施地面安全感知中的有效性。
  • SONG Dongping
    Accepted: 2026-07-31
    The low-altitude economy has been incorporated into the national strategic emerging industries. The low-altitude Intelligent Integrated Electronic System (LIIES) serves as the core technological framework ensuring the safe and efficient operation of low-altitude aircraft. It is evolving from a fragmented architecture towards an integrated, intelligent, and networked direction.. This paper provides a systematic exposition of the system's fundamental concepts, architecture, and key components, and presents a comparative analysis against conventional avionics systems used in large transport aircraft. The state of the art, application bottlenecks, and future trends are examined across several critical dimensions, including onboard communication, navigation, and surveillance (CNS), flight management, multi-modal human-machine interaction, detect-and-avoid (DAA) capabilities, low-altitude intelligent connectivity, and ground-based support infrastructure. The paper concludes by summarizing the primary technical challenges and outlining prospective development pathways centered on standardization, modularization, and enhanced system autonomy. Collectively, these efforts aim to provide both a theoretical foundation and technical reference for the large-scale commercialization of the low-altitude economy.
  • 张 清 秀, 何 畅, 王 君 晓
    Accepted: 2026-07-27
    To address the issues of high-frequency calls, large computational load, and dynamic parameter changes with subframes in the Linear Predictive Coding (LPC) analysis filter of the Opus audio encoder's SILK mode, this paper designs and implements a reconfigurable hardware accelerator based on the Zynq-7000 SoC platform. This design employs a parameterized transposed FIR fully pipelined architecture, combined with DSP48E1 resource optimization, achieving one input sample per clock cycle in steady state. To ensure data consistency during variable-length subframe switching, a Shadow/Active double-buffered parameter update mechanism is introduced, ensuring filter parameters are committed at processing frame boundaries and preventing parameter tearing within the same processing frame. An efficient hardware-software co-processing data path is constructed using AXI DMA, and hardware implementation and verification are completed according to Opus fixed-point arithmetic rules. Experimental results show that the LUT utilization rate of this design is less than 2.1% on the XC7Z100 platform, and it meets timing constraints at the target clock frequency. Under typical wideband long frame configurations, compared with the optimized libopus fixed-point C reference implementation running on ARM Cortex-A9, the algorithm core speedup reaches 9.98 times, and the average end-to-end speedup reaches 3.82×. The hardware output maintains bit-exact consistency with the software reference model. This design has low resource overhead while ensuring computational correctness, and can provide an effective hardware acceleration solution for resource-constrained embedded audio coding systems.
  • Ma Jie, Zhao Zhengjie, Ning Sihua
    Accepted: 2026-07-20
    To address the timing closure challenges in modern Very-Large-Scale Integration (VLSI) design and the limitations of existing Local Clock Buffer (LCB) optimization and Lagrangian Relaxation (LR) methods, which often fall into local optima in large-scale designs, this paper proposes a collaborative optimization framework. The framework integrates coarse-grained timing balancing with LCB reallocation and fine-grained relative placement strategies. At the global level, heuristic algorithms are employed to achieve fast optimization, while at the local level, refined search techniques further enhance timing quality, thus achieving an effective trade-off between placement constraints and timing optimization. Experimental results on eight ICCAD-2015 benchmark circuits demonstrate that the proposed method consistently outperforms state-of-the-art approaches: in setup timing optimization, the Worst Negative Slack (WNS) and Total Negative Slack (TNS) are improved by 2% and 4%, respectively; in hold timing optimization, WNS and TNS are improved by 14% and 1%, respectively. These results validate the effectiveness and superiority of the proposed framework in improving both timing performance and optimization efficiency.
  • Accepted: 2026-07-20
    目前多视频多目标的跟踪识别是大多都是通过单独的嵌入式RK3588型号SOC 系统实现,但是整个系统明显存在接口扩展性差,运行大型复杂程序硬件资源不足的问题。本文设计了一套基于国产化的FPGA作为前端接口板采集多路视频或者其他传感器信号,以Atlas 200I 作为AI板进行识别算法加速推理,以RK3588作为CPU运行复杂的显控程序,构建了一套基于CPU板、AI板和接口板的硬件框架。整个系统可以实现红外和可见光的Cameralink和SDI接口视频输入,以及其他传感器信号输入。针对于该系统在基于CPU板的嵌入式国产的麒麟系统上设计一套多路视频识别与跟踪显控系统。CPU板的显控系统能监控各个板卡硬件的基本参数,以及实现2路视频跟踪标记以及存储。实验结果表明,整个系统能够支持高速多路的边缘推理,显控软件满足多路视频识别与跟踪显示与画框标记,该系统具有一定的实际意义。
  • Song Yaowei
    Accepted: 2026-07-14
    Abstract: This paper proposes a Configurable Ring Oscillator PUF driven by a Grouped Inter-stage Exchange Network (GIEN-CROPUF). Adopting the approach of grouping + intra-group permutation + a small number of inter-group exchanges, this structure continuously restructures the RO links between adjacent stages, thereby achieving a higher complexity of path combinations and effectively destroying the decomposability of the delay model of traditional CROPUFs. Comparative experiments were conducted based on a variety of machine learning modeling attacks (e.g., MLP, SVM, Logistic Regression). The results show that within the scale of feasible Challenge-Response Pairs (CRPs), the prediction accuracy of the proposed structure is significantly lower than that of traditional structures.
  • Sun Xueli
    Accepted: 2026-07-14
    This paper addresses the issues of high false detection rates due to environmental interference, response lag in traditional tracking methods, and insufficient aiming accuracy in mobile target recognition. A fast visual tracking and aiming system based on machine vision is designed and implemented. The system employs the TI MSP0G3507 microcontroller as the core, constructing a hardware platform composed of a two-dimensional stepper motor gimbal, an OpenMV-H7 Plus vision module, a laser pointer, and a four-wheel line-following chassis. In terms of visual processing, a composite recognition algorithm integrating geometric features, pixel density, and spatial relational constraints is proposed, significantly enhancing the recognition robustness of A4 paper rectangular targets in complex backgrounds. For the two-dimensional gimbal control, digital PID control technology is adopted to convert image coordinate deviations into gimbal control inputs, achieving precise aiming. Regarding platform path tracking, a multi-channel grayscale sensor combined with a PID line-following algorithm enables stable path tracking. Experimental results show that in an indoor environment, the static aiming error of the system is less than 2 mm, the tracking delay for moving targets at 0.5 m/s is below 200 ms, and the deviation of a laser-drawn circular trajectory with a radius of 10 cm is less than 2 mm. The system features high integration, compact size, and realizes the coordinated operation of recognition, aiming, and movement, providing a feasible engineering solution for visual servoing applications on embedded mobile platforms.
  • Accepted: 2026-07-02
    This paper describes a LDO with low 1/f noise and wide bandwidth. With the help of chopping technology, the achievable output noise of LDO is 139nv/√Hz at 10Hz, 64.7nv/√Hz at 100Hz, and 36.3nv/√Hz at 1kHz under the worst corner. Since the 1/f noise corner is as low as Hertz level, it can be used to power the modules whose signal bandwidth is Hertz level. The LDO proposed in this paper also adds the current fast compensation mechanism (CFA) and the fast compensation current mechanism at the load end of the output regulator, which can not only achieve a very high SNR, but also improve the power supply suppression energy and fast power supply capacity. The CFA also splits the low frequency pole between the amplifier and the PMOS regulator into two higher frequency poles, which extends the LDO bandwidth and achieves a higher PSRR.
  • Accepted: 2026-07-02
    In response to the problems of delayed response in elderly care at home and in the community, inconvenient positioning, inability to call for help in case of sudden illness, and easy getting lost, this paper designs an intelligent monitoring system suitable for the elderly. The system takes the STC12C5A60S2 single-chip microcomputer as the core, integrates the PulseSensor heart rate sensor, DS18B20 temperature sensor, GT-U7 GPS module and buzzer, and adopts the "hardware integration + software control" mode. Data transmission is achieved through single-wire and serial communication without the need for additional conversion modules, simplifying the hardware design. The system can accurately monitor the physiological parameters of the elderly such as heart rate and body temperature, output real-time positioning information, and automatically trigger an alarm in case of abnormality. It operates stably and is easy to operate. This system effectively compensates for the shortcomings of traditional monitoring, can comprehensively ensure the safety of the elderly, and has controllable costs and strong practicality. It is suitable for home and community elderly care scenarios and can meet the daily monitoring needs of the elderly, with high promotion value and application prospects.
  • WU zhou Xiong, DAI di Xiao, HE quan Yu
    Accepted: 2026-07-02
    In order to improving the storage efficiency of NVMe SSD in airborne embedded storage system under complex task scenarios and ensuring the operational stability and accuracy of the embedded storage system, this paper does some research on reading-writing architecture of NVMe SSD based on synchronous multi-channel technology. By elaborating the design methods of I/O queues, physical region pages and interrupt processing under the synchronous multi-channel architecture, the implementation process and operating characteristics of the synchronous multi-channel NVMe SSD read-write architecture are intuitively presented. Meanwhile, a real test scenario is constructed based on a multi-core processor platform and deployed in an actual airborne embedded storage system. The experimental results verify the correctness of the proposed synchronous multi-channel NVMe SSD read-write architecture.
  • YOU Yong, LUO Bingyin
    Accepted: 2026-07-02
    In power converters, extended soft-start duration and loop compensation typically require large capacitors. However, off-chip capacitors increase packaging cost and peripheral complexity, reducing system reliability. Although on-chip integration is viable, capacitance density is strictly limited by process constraints and chip area, thereby restricting compensation network flexibility and soft-start time adjustment. To address these limitations, a highly integrated DC-DC converter with a simplified off-chip environment is proposed. For internal loop compensation, an improved scheme employing a transconductance amplifier achieves equivalent capacitance amplification, enhancing stability and design flexibility. For on-chip soft-start, a capacitor-free method utilizing charging current diversion is presented, enabling prolonged soft-start time within a limited area to suppress inrush current effectively. The converter is implemented in a 0.18-µm BCD process with an area of 1.73 mm². Measurement results demonstrate that, at a 5 V/2 A output rating, the converter achieves an efficiency exceeding 80% over a load range of 1 mA to 2 A, with a peak efficiency of 95.1%. Under a load step change, the output voltage overshoot and undershoot are less than 180 mV, and the recovery time is approximately 200 μs, indicating good stability of the converter. The measured start up time is approximately 4.8 ms.
  • chen qin jia, shao feng
    Accepted: 2026-07-01
    At the 180 nm CMOS process node, multi-objective optimization of standard-cell libraries still relies heavily on SPICE simulation, resulting in severe computational bottlenecks. A single transient analysis typically takes ~100 ms; a typical library optimization requiring over 10 000 iterations consumes approximately 2.1 minutes even with 8-core parallelization. This paper proposes a lightweight surrogate modeling framework that constructs a hybrid dataset of only 155 samples by fusing 25 high-fidelity PySpice simulations (GF180MCU PDK) with 130 open-source benchmark points. After systematically comparing eight machine-learning architectures, the multi-layer perceptron (MLP) model achieves the best accuracy-speed trade-off (R² = 0.8734 for power and 0.9521 for delay) with an inference time of only 1μs, delivering approximately 1 00 000× acceleration over conventional SPICE. Zero-shot cross-cell generalization yields an average R² of 0.6746; fine-tuning with merely 15 additional SPICE samples per new cell further improves performance. Industrial ROI analysis for a 50-engineer design team shows the design cycle reduced from 2.1 min to 0.5 min, with a 3-year net benefit yielding ROI of 200–300 %. The proposed approach provides a deployable, cost-effective solution for AI-assisted circuit design automation with substantial engineering and economic value.
  • zhang xiaowen
    Accepted: 2026-06-29
    Targeting the characteristics of GJB/Z 299C-2006 "Reliability Prediction Handbook for Electronic Equipment," where predicted reliability data for various components, including integrated circuit chips under operational conditions, often deviate significantly from actual situations, a failure mechanism-based reliability data validation method has been designed. This validation method starts from the failure mechanisms in very large-scale integrated circuit chips and, based on an analysis of domestic and international standards for evaluating reliability related to failure mechanisms, summarizes the failure physics models in national standards. Leveraging reliability test structures, through accelerated life testing for single failure mechanisms, the median time to failure for primary failure mechanisms is derived. By converting the median time to failure into failure rates, reliability data for very large-scale integrated circuits is obtained, ensuring the validity of the reliability data from a physical mechanism. These reliability data can be applied to the reliable utilization of avionics products while also supporting reliability prediction work for electronic products.
  • Accepted: 2026-06-29
    Network-on-Chip (NoC) has been widely adopted for on-chip communication due to its advantages of high integration, strong scalability, and low power consumption. This paper designs and implements a highly reliable dual-layer Network-on-Chip circuit. A redundant architecture is adopted via two NoC layers: one layer is used for normal data transmission, while the other layer is dedicated to timeout retransmission. To enhance anti-interference capability, triple modular redundancy (TMR) is applied to key routing control information of event packets, including flit type coding, destination router node information, and NoC ID. Meanwhile, parity check is used for the carried address and data information to ensure the correctness of data transmission. Simulation results verify that the proposed highly reliable dual-layer NoC circuit achieves correct functional behavior and excellent fault-tolerant performance.
  • Accepted: 2026-06-26
    With the widespread application of FPGAs in high-performance embedded computing and data centers, the demand for data transmission bandwidth via the PCIe interface is increasing. Xilinx's XDMA IP core, as a mainstream high-performance DMA solution, often has its actual performance limited by the complex memory management mechanism of the Linux system. This paper studies the key points affecting the XDMA transmission performance in the standard Linux driver model through theoretical and modeling analysis, and finds that the "lazy allocation" strategy of user-space memory causes the allocation and mapping of physical pages to be delayed until after the DMA transmission request is initiated, frequently triggering page faults and increasing the TLB miss rate, which seriously restricts the efficiency and determinacy of high-bandwidth transmission. This paper proposes an application layer memory pre-mapping optimization strategy that utilizes advanced parameters of the mmap system call. This strategy moves the physical memory allocation, page table establishment, and page locking operations forward to the system initialization stage, thereby reducing the runtime overhead and significantly improving the subsequent memory access efficiency. Theoretical analysis and experimental results show that this strategy increases the data transmission rate by 85.5% under the default TLB size compared to the optimized version. Furthermore, the impact of TLB size on XDMA transmission is studied, which is of great reference significance for building high-performance, low-latency embedded heterogeneous systems.
  • SHI Huanhuan, LI Sujuan, BAO Zhong
    Accepted: 2026-06-17
    Time sensitive network has the characteristics of low latency, low cost, and high reliability. The Time-Aware Shaper (TAS), defined by the IEEE 802.1Qbv standard, stands as one of the critical technologies for enabling deterministic network traffic and holds a core position within the TSN protocol suite. First, by means of logical designs such as top-level architecture, workflow, and active-standby coordination, a method for implementing TAS is proposed. Second, by constructing constraints for traffic with different priorities, a deterministic scenario is designed to verify the implementation of clock synchronization and the time-aware shaper. Finally, through comparative experiments conducted before and after enabling TAS, the superiority of TSN technology over standard Ethernet in terms of low latency and low jitter is validated. The results show that, after enabling clock synchronization and gate control functions, the average delay and jitter are only about 1/20 to 1/15 of the original, which is significantly different. This provides a demonstration solution for the rapid promotion of TSN technology.
  • SU He, TANG Wei, WANG yi Jing, CHEN Lin
    Accepted: 2026-06-17
    Addressing the power supply issues of noise-sensitive electronic devices such as mobile phone cameras and Bluetooth, a low dropout regulator (LDO) with high power supply rejection ratio (PSRR) and ultra-low dropout voltage has been designed. The circuit employs an N-type field-effect transistor (NMOS) as the regulation transistor and is powered by a dual-power supply. It segregates the bias power supply from the input power supply of the regulation transistor, thereby attaining a high Power Supply Rejection Ratio (PSRR) and an ultra-low dropout voltage. The design incorporates a pre-regulation modulation circuit and a low-pass filter to process the reference voltage, enhancing the PSRR of the bias power supply. The main loop adjust the system's pole-zero distribution through inverse nested Miller compensation, improving the overall PSRR of the circuit. The circuit and layout design were completed based on a 0.18µm CMOS process. The maximum load current of the circuit is 500mA. Simulation results show that the dropout voltage at maximum load current is 100mV. At a load of 10mA, the power supply rejection ratio (PSRR) of the input power supply at frequencies of 100Hz, 1kHz, 10kHz, and 1MHz are -110dB, -90dB, -70dB, and -65dB, respectively.
  • 张 剑
    Accepted: 2026-06-17
    Spiking neural networks (SNNs) represent and process information using discrete spike trains, exhibiting event-driven and sparse computation characteristics and offering strong potential for energy-efficient intelligent computing. To fully exploit their low-power advantages, dedicated hardware implementation techniques and neuromorphic chips for SNNs are of significant research importance. From the perspective of hardware implementation, this paper systematically reviews the full hardware development chain of SNNs, covering neuron models, information encoding methods, network structures, and hardware architectures. The paper describes commonly used spiking neuron models and information encoding methods, discusses network structures including fully connected, convolutional, and attention-based architectures, systematically summarizes research on hardware architectures, compares the advantages and limitations of different circuit implementation technologies and computing paradigms, and analyzes current key challenges while outlining future research directions.