Home Browse Just Accepted

Just Accepted

Note: The articles listed below have been peer-reviewed and accepted for publication in this journal. These articles have not yet been scheduled for a specific issue; their content and layout may undergo minor changes in the final published version. Please refer to the final published version as the definitive one. This journal has assigned each of these articles a unique and persistent DOI. You may use the DOI to cite this article directly.
Please wait a minute...
  • Select all
    |
  • Accepted: 2026-09-21
    For the 14nm FinFET advanced process node, traditional BSIM physical models have long calibration cycles and poor adaptability to new devices, lookup table models have insufficient accuracy, and machine learning device models are difficult to access commercial EDA simulation tools. This paper proposes a complete two-layer collaborative DTCO optimization process of ANN device proxy + random forest circuit proxy. First, a fully connected artificial neural network with batch normalization and LeakyReLU activation is built based on the Sentaurus TCAD multi-dimensional simulation dataset, integrating gate structure parameters, aging stress, and electrical bias inputs to accurately predict FinFET leakage current and terminal charges. An automated parsing and conversion tool is designed to export the trained PyTorch network into a Verilog-A device model compatible with Spectre/HSPICE, connecting the machine learning model with the SPICE simulation link. Latin hypercube sampling is used to traverse the high-dimensional design space of device-circuit coupling, and after batch simulation, a random forest lightweight circuit proxy model is trained to replace time-consuming repeated circuit simulations. Based on the NSGA-III multi-objective evolutionary algorithm, the Pareto optimal solution set is solved with the full-adder unit delay and power consumption as optimization objectives. Experimental results show that the test set determination coefficient of the FinFET device ANN model is R2=0.996, which is 105 times faster than the native TCAD simulation; the overall acceleration of the hierarchical proxy model collaborative optimization process is 30 times; the optimized design can achieve a 22% reduction in delay under the same power consumption and a 31% reduction in power consumption under the same delay. The proposed process can fully realize cross-layer joint optimization of device-process-circuit including reliability aging variables, providing a standardized technical solution for the automated implementation of advanced process DTCO automation.
  • Accepted: 2026-09-17
    目前,自动测试设备(ATE)的测试通道数量限制了单次可测试的集成芯片(IC)数量。本文提出一种基于电容耦合的多IC并行无线测试架构。该架构将测试板电容极板上的信号耦合至IC底部的电容极板,信号经硅通孔(TSV)传输至中介层中的测试电路。通过系统信息块(SIB)控制信号进入小芯片的路径,利用电容耦合实现测试信号的无线传输。本文通过仿真实验验证了所提方法的可行性。结果表明,该方法能够实现并行无线测试的可行性并具备潜在的并行测试能力,无线测试信号传输速率可达119.21 MB/s。
  • wang mingyu
    Accepted: 2026-09-09
    To address the frequency-dependent rank growth and the resulting loss of compression efficiency when adaptive cross approximation (ACA) is directly applied to the loop-impedance matrix in broadband magneto-quasi-static resistance and inductance (RL) extraction, a basis-transformation-based ACA scheme is proposed. The directional degrees of freedom of vector current basis functions are isolated into three sparse mapping matrices, leaving a dense core matrix that contains only scalar couplings between panels. ACA is then applied exclusively to this core matrix. The formulation is combined with mixed surface meshes, loop-current analysis, and a broadband equivalent surface-impedance model. Numerical results show that the rank and compression ratio of far-field blocks remain nearly frequency independent. For the two-conductor example, the maximum speedups of matrix filling and compression are approximately 6 and 80, respectively. For the FCBGA and FCCSP examples, matrix-generation time is reduced from 89 s to 13 s and from 304 s to 25 s, while peak memory is reduced by 20.8% and 31.3%. The maximum reported deviations of the extracted R and L parameters are 4.6% and 1.73%, respectively.
  • Accepted: 2026-09-09
    To meet the demands of humanoid robot joint actuation for high power density, high efficiency, fast response, and high reliability, this paper systematically reviews the key technologies and development trends of motor gate driver ICs. It first analyzes the differentiated requirements of various joints in terms of power, voltage, and control performance, and compares the technical characteristics of silicon-based and GaN power devices. Then, focusing on four major trade-offs—switching speed versus electromagnetic interference, dead-time optimization, signal transmission rate versus common-mode transient immunity, and power density versus thermal integration—it surveys promising technologies including active gate driving, dead-time optimization, high-reliability signal transmission, monolithic integration, and intelligent sensing. Finally, it summarizes the technology roadmaps for different joint types and points out that high-frequency operation, integration, intelligence, and specialization will be the future development directions.
  • Accepted: 2026-09-08
    A wide-dynamic-range reconfigurable logarithmic amplifier based on 0.18 μm CMOS process is designed to address the problems of wide bandwidth of partial discharge signals of electrical equipment and significant amplitude difference in sensor output signals. To balance broadband response and wide dynamic range detection capability, a successive detection logarithmic amplification architecture is adopted, which combines multi-stage limiting amplification with full-wave detection circuits to realize logarithmic compression of input signals. To improve bias stability and reduce the influences of process and temperature variations, a constant-transconductance self-bias circuit is introduced to enhance the consistency of bias current and static operating point of each limiting amplification unit. Aiming at the differentiated requirements of logarithmic slopes in different detection scenarios, a logarithmic response slope reconfiguration method based on an on-chip digital switch array is proposed. With the main circuit structure unchanged, flexible configuration of multiple logarithmic response slopes is realized by selecting outputs of different detection current branches through the digital switch array. Simulation results show that the designed amplifier achieves a −3 dB bandwidth of 40 MHz and an effective input dynamic range of 72 dB under a 2 V supply voltage, with the logarithmic conformity error within ±1.9 dB and eight adjustable slope gears supported. Integrating broadband performance, wide dynamic range and reconfigurable characteristics, the proposed amplifier can meet the application requirements of partial discharge online monitoring systems for multi-sensor adaptation, wide-dynamic-range detection and multi-gear output configuration.
  • Liangshun Wu
    Accepted: 2026-09-03
    大语言模型(Large Language Models,LLMs)的快速发展,为硬件/软件协同设计中的代码自动生成带来了新的可能。本文提出一种面向多处理器片上系统(MPSoC)平台的AMBA3-AXI互连代码生成混合框架,将基于模板的确定性生成方法与提示工程相结合。该框架由参数列表生成器、模板化代码生成器和LLM驱动的智能代码生成器三部分组成,在保持代码结构可预测性的同时,提升了对动态设计需求的适应能力。基于Specman + eVC验证环境的综合评估表明,该框架能够提升可扩展性、减少人工工作量并提高代码正确性。实验结果进一步证明,引入LLM有助于改善代码可维护性、增强协议一致性,并缩短MPSoC设计与验证周期。
  • Accepted: 2026-08-17
    To address the engineering challenges of fragmented toolchains and inefficient data flow in semiconductor device modeling and Design Technology Co-Optimization (DTCO) workflows, this paper presents the design and implementation of SCNN (Smartchip Neural Network), an integrated software platform. By leveraging established machine learning and multi-objective optimization algorithms, SCNN adopts a hierarchical modeling strategy to construct an automated DTCO workflow spanning from device-level data processing to circuit-level design space exploration. The platform integrates 13 machine learning algorithms and 25 multi-objective evolutionary algorithms, supporting six core functions: device simulation, data processing, device predictive model training, Verilog-A compatible neural network model export, circuit predictive model training, and multi-objective optimization. In MOSFET DC characteristic modeling experiments, the algorithms integrated within SCNN consistently achieve a coefficient of determination R² > 0.995 fitting accuracy. The platform enables automated conversion of trained neural networks into Verilog-A compatible models, which are suitable for DC operating point analysis and low-frequency AC simulation scenarios. A multi-objective optimization case study of an operational amplifier demonstrates the effectiveness of the hierarchical modeling strategy for efficient circuit design space exploration. The platform has been validated in multiple industrial projects.
  • Ren Zuowei, Yuan Yiming, Feng Min, Dai Ning, Jiao Yang
    Accepted: 2026-08-13
    Existing DDS waveform generators cannot simultaneously suppress phase noise and spurs, failing to balance time- and frequency-domain performance of arbitrary waveforms. This paper builds an FPGA+AD9910 waveform generation system. Isolated LDO power supplies and low-jitter clocks suppress noise at source; the mechanisms of spurs and phase noise are analyzed with a spectral error model built. A 10 MHz 7th-order Bessel low-pass filter is designed to suppress spurs and output distortion-free waveforms.Simulations and measurements show the filter’s passband group delay ripple <1% and 125 dB stopband attenuation at 50 MHz. The system achieves 50.57 dB spurious rejection, one-order lower THD for arbitrary waveforms, and 33 dB phase noise improvement at 100 Hz offset of a 100 MHz carrier, satisfying radar and precision test demands.
  • YANG Le, TANG Wei, HE Ming-liang, XU Xiao-ning, HE Zheng-rong
    Accepted: 2026-08-04
    To address the issue that traditional bandgap reference circuits tend to introduce additional power consumption when reducing temperature drift, a low-power and low-temperature drift bandgap reference source with nA-level static current has been designed. This circuit employs a clamp structure without the traditional bias branch to generate PTAT current, avoiding the additional bias current caused by traditional operational amplifier clamping; at the same time, a segmented curvature compensation module based on the low bias transconductance comparison branch is introduced to generate compensation currents in the low-temperature and high-temperature sections respectively, in order to suppress the high-order temperature curvature in the first-order compensated reference voltage. The key MOS transistors in the compensation branch operate in the sub-threshold region, and the exponential characteristic of sub-threshold current is utilized to achieve low-power temperature compensation. Additionally, a transistor collector-to-substrate leakage current compensation circuit is introduced to reduce the influence of high-temperature leakage current on the reference output voltage. Simulation results show that at a 3.3 V power supply voltage and within the temperature range of -40 to 125 ℃, the bandgap reference output voltage is 1.209 V, the temperature coefficient is 2.33 ppm/℃, the low-frequency power supply rejection ratio is 64 dB, the static current is 365 nA, and the corresponding static power consumption is approximately 1.20 μW.
  • Feng Xuning, Zhang Deming, Hu Yuanqi, Zhu Dapeng, Cheng Yuanqing
    Accepted: 2026-08-03
    Angular encoder chips often operate in complex environments and impose strict requirements on signal filtering. However, constrained by limited on-chip space, integrating complex filtering circuits remains a significant challenge. To address this issue, this article proposes a folded Kalman filter circuit that can co-optimize area and latency. The design employs four optimization strategies—a folded architecture, filter model splitting, Kalman parameter pruning, and operation sequence reordering—which significantly reduce hardware overhead. Processing a single-channel signal requires only 2 adders, 1 multiplier, and 1 divider. Implemented in 180nm process node at a 50MHz clock frequency, the circuit supports high-speed signal processing with a sampling rate of up to 2MSPS, achieving high-frequency noise attenuation up to 22dB. Without compromising filtering accuracy, our design reduces DSP resource consumption by 68% and computational latency by 53% compared to related works. It can be effectively applied to miniature signal processing chips, such as angular encoders.
  • Accepted: 2026-08-03
    【目的】目前如光传输网络OTN(Optical Transmission Network)、以太网数据通信传输的研究热点都集中在高速率传输,例如40Gbps以上,相应的串行器/解串器SerDes(Serializer-Deserializer)技术和芯片也随之推出,这些高速SerDes都存在下限速率,目前SerDes芯片存在速率下限,通常在500Mbps左右,因此500Mbps以下的数据串行传输不能直接使用SerDes,为了解决此问题,本文深入分析了过采样与SerDes结合的数据传输,并重点研究了接收端的过采样数据恢复。【方法】本文给出了基于SerDes的时域过采样设计,并详细设计了接收端的过采样数据恢复模块,给出了设计接口和信号说明;针对发送端、接收端的时钟频差时的过采样给出了详细分析和采样补偿或者舍弃,而且讨论了长连1或者长连0的最大容许长度,以便选择何种物理层编码码型。最后,对两种频差情况进行了详细仿真和讨论,并且基于紫光国产FPGA芯片PG2T100的板卡开展了测试验证。【结果】仿真和测试结果表明,基于SerDes的时域过采样设计系统满足500Mbps以下的数据传输需求。【结论】基于SerDes的时域过采样系统设计详细、仿真全面、上板测试充分,为500Mbps以下的数据串行传输提供了工程设计指导。
  • 张 清 秀, 何 畅, 王 君 晓
    Accepted: 2026-07-27
    To address the issues of high-frequency calls, large computational load, and dynamic parameter changes with subframes in the Linear Predictive Coding (LPC) analysis filter of the Opus audio encoder's SILK mode, this paper designs and implements a reconfigurable hardware accelerator based on the Zynq-7000 SoC platform. This design employs a parameterized transposed FIR fully pipelined architecture, combined with DSP48E1 resource optimization, achieving one input sample per clock cycle in steady state. To ensure data consistency during variable-length subframe switching, a Shadow/Active double-buffered parameter update mechanism is introduced, ensuring filter parameters are committed at processing frame boundaries and preventing parameter tearing within the same processing frame. An efficient hardware-software co-processing data path is constructed using AXI DMA, and hardware implementation and verification are completed according to Opus fixed-point arithmetic rules. Experimental results show that the LUT utilization rate of this design is less than 2.1% on the XC7Z100 platform, and it meets timing constraints at the target clock frequency. Under typical wideband long frame configurations, compared with the optimized libopus fixed-point C reference implementation running on ARM Cortex-A9, the algorithm core speedup reaches 9.98 times, and the average end-to-end speedup reaches 3.82×. The hardware output maintains bit-exact consistency with the software reference model. This design has low resource overhead while ensuring computational correctness, and can provide an effective hardware acceleration solution for resource-constrained embedded audio coding systems.
  • Ma Jie, Zhao Zhengjie, Ning Sihua
    Accepted: 2026-07-20
    To address the timing closure challenges in modern Very-Large-Scale Integration (VLSI) design and the limitations of existing Local Clock Buffer (LCB) optimization and Lagrangian Relaxation (LR) methods, which often fall into local optima in large-scale designs, this paper proposes a collaborative optimization framework. The framework integrates coarse-grained timing balancing with LCB reallocation and fine-grained relative placement strategies. At the global level, heuristic algorithms are employed to achieve fast optimization, while at the local level, refined search techniques further enhance timing quality, thus achieving an effective trade-off between placement constraints and timing optimization. Experimental results on eight ICCAD-2015 benchmark circuits demonstrate that the proposed method consistently outperforms state-of-the-art approaches: in setup timing optimization, the Worst Negative Slack (WNS) and Total Negative Slack (TNS) are improved by 2% and 4%, respectively; in hold timing optimization, WNS and TNS are improved by 14% and 1%, respectively. These results validate the effectiveness and superiority of the proposed framework in improving both timing performance and optimization efficiency.
  • Accepted: 2026-07-20
    目前多视频多目标的跟踪识别是大多都是通过单独的嵌入式RK3588型号SOC 系统实现,但是整个系统明显存在接口扩展性差,运行大型复杂程序硬件资源不足的问题。本文设计了一套基于国产化的FPGA作为前端接口板采集多路视频或者其他传感器信号,以Atlas 200I 作为AI板进行识别算法加速推理,以RK3588作为CPU运行复杂的显控程序,构建了一套基于CPU板、AI板和接口板的硬件框架。整个系统可以实现红外和可见光的Cameralink和SDI接口视频输入,以及其他传感器信号输入。针对于该系统在基于CPU板的嵌入式国产的麒麟系统上设计一套多路视频识别与跟踪显控系统。CPU板的显控系统能监控各个板卡硬件的基本参数,以及实现2路视频跟踪标记以及存储。实验结果表明,整个系统能够支持高速多路的边缘推理,显控软件满足多路视频识别与跟踪显示与画框标记,该系统具有一定的实际意义。
  • Song Yaowei
    Accepted: 2026-07-14
    Abstract: This paper proposes a Configurable Ring Oscillator PUF driven by a Grouped Inter-stage Exchange Network (GIEN-CROPUF). Adopting the approach of grouping + intra-group permutation + a small number of inter-group exchanges, this structure continuously restructures the RO links between adjacent stages, thereby achieving a higher complexity of path combinations and effectively destroying the decomposability of the delay model of traditional CROPUFs. Comparative experiments were conducted based on a variety of machine learning modeling attacks (e.g., MLP, SVM, Logistic Regression). The results show that within the scale of feasible Challenge-Response Pairs (CRPs), the prediction accuracy of the proposed structure is significantly lower than that of traditional structures.
  • Sun Xueli
    Accepted: 2026-07-14
    This paper addresses the issues of high false detection rates due to environmental interference, response lag in traditional tracking methods, and insufficient aiming accuracy in mobile target recognition. A fast visual tracking and aiming system based on machine vision is designed and implemented. The system employs the TI MSP0G3507 microcontroller as the core, constructing a hardware platform composed of a two-dimensional stepper motor gimbal, an OpenMV-H7 Plus vision module, a laser pointer, and a four-wheel line-following chassis. In terms of visual processing, a composite recognition algorithm integrating geometric features, pixel density, and spatial relational constraints is proposed, significantly enhancing the recognition robustness of A4 paper rectangular targets in complex backgrounds. For the two-dimensional gimbal control, digital PID control technology is adopted to convert image coordinate deviations into gimbal control inputs, achieving precise aiming. Regarding platform path tracking, a multi-channel grayscale sensor combined with a PID line-following algorithm enables stable path tracking. Experimental results show that in an indoor environment, the static aiming error of the system is less than 2 mm, the tracking delay for moving targets at 0.5 m/s is below 200 ms, and the deviation of a laser-drawn circular trajectory with a radius of 10 cm is less than 2 mm. The system features high integration, compact size, and realizes the coordinated operation of recognition, aiming, and movement, providing a feasible engineering solution for visual servoing applications on embedded mobile platforms.
  • Accepted: 2026-07-02
    This paper describes a LDO with low 1/f noise and wide bandwidth. With the help of chopping technology, the achievable output noise of LDO is 139nv/√Hz at 10Hz, 64.7nv/√Hz at 100Hz, and 36.3nv/√Hz at 1kHz under the worst corner. Since the 1/f noise corner is as low as Hertz level, it can be used to power the modules whose signal bandwidth is Hertz level. The LDO proposed in this paper also adds the current fast compensation mechanism (CFA) and the fast compensation current mechanism at the load end of the output regulator, which can not only achieve a very high SNR, but also improve the power supply suppression energy and fast power supply capacity. The CFA also splits the low frequency pole between the amplifier and the PMOS regulator into two higher frequency poles, which extends the LDO bandwidth and achieves a higher PSRR.
  • Accepted: 2026-07-02
    In response to the problems of delayed response in elderly care at home and in the community, inconvenient positioning, inability to call for help in case of sudden illness, and easy getting lost, this paper designs an intelligent monitoring system suitable for the elderly. The system takes the STC12C5A60S2 single-chip microcomputer as the core, integrates the PulseSensor heart rate sensor, DS18B20 temperature sensor, GT-U7 GPS module and buzzer, and adopts the "hardware integration + software control" mode. Data transmission is achieved through single-wire and serial communication without the need for additional conversion modules, simplifying the hardware design. The system can accurately monitor the physiological parameters of the elderly such as heart rate and body temperature, output real-time positioning information, and automatically trigger an alarm in case of abnormality. It operates stably and is easy to operate. This system effectively compensates for the shortcomings of traditional monitoring, can comprehensively ensure the safety of the elderly, and has controllable costs and strong practicality. It is suitable for home and community elderly care scenarios and can meet the daily monitoring needs of the elderly, with high promotion value and application prospects.
  • WU zhou Xiong, DAI di Xiao, HE quan Yu
    Accepted: 2026-07-02
    In order to improving the storage efficiency of NVMe SSD in airborne embedded storage system under complex task scenarios and ensuring the operational stability and accuracy of the embedded storage system, this paper does some research on reading-writing architecture of NVMe SSD based on synchronous multi-channel technology. By elaborating the design methods of I/O queues, physical region pages and interrupt processing under the synchronous multi-channel architecture, the implementation process and operating characteristics of the synchronous multi-channel NVMe SSD read-write architecture are intuitively presented. Meanwhile, a real test scenario is constructed based on a multi-core processor platform and deployed in an actual airborne embedded storage system. The experimental results verify the correctness of the proposed synchronous multi-channel NVMe SSD read-write architecture.
  • YOU Yong, LUO Bingyin
    Accepted: 2026-07-02
    In power converters, extended soft-start duration and loop compensation typically require large capacitors. However, off-chip capacitors increase packaging cost and peripheral complexity, reducing system reliability. Although on-chip integration is viable, capacitance density is strictly limited by process constraints and chip area, thereby restricting compensation network flexibility and soft-start time adjustment. To address these limitations, a highly integrated DC-DC converter with a simplified off-chip environment is proposed. For internal loop compensation, an improved scheme employing a transconductance amplifier achieves equivalent capacitance amplification, enhancing stability and design flexibility. For on-chip soft-start, a capacitor-free method utilizing charging current diversion is presented, enabling prolonged soft-start time within a limited area to suppress inrush current effectively. The converter is implemented in a 0.18-µm BCD process with an area of 1.73 mm². Measurement results demonstrate that, at a 5 V/2 A output rating, the converter achieves an efficiency exceeding 80% over a load range of 1 mA to 2 A, with a peak efficiency of 95.1%. Under a load step change, the output voltage overshoot and undershoot are less than 180 mV, and the recovery time is approximately 200 μs, indicating good stability of the converter. The measured start up time is approximately 4.8 ms.