A Chinese AI chip combining software-defined computing with 3D near-memory architecture delivered 520 trillion floating-point operations per second (TFLOPS) on a 14nm manufacturing process, debuting in Shanghai on July 131,2.
The chip is described as China's first to pair these two design approaches, achieving high throughput without relying on leading-edge lithography. Rather than pursuing smaller transistor geometries, the design extracts performance through architectural choices that address two persistent bottlenecks: rigid hardware utilization and the memory wall.
The software-defined architecture allows hardware resources to be dynamically reconfigured for different computing tasks, raising compute utilization rates relative to fixed-function accelerators. The chip can handle workloads spanning image recognition, natural language processing, and scientific computing without the efficiency penalties typically associated with general-purpose hardware.
A full-stack software toolchain compatible with mainstream deep learning frameworks shipped alongside the silicon. The product portfolio spans single accelerator cards, AI servers, liquid-cooled supernodes, and large-scale intelligent computing clusters, indicating the architecture is designed to scale from edge devices to data-center-scale deployments.
The 3D near-memory computing approach vertically stacks compute units with memory, achieving 6.4 TB per second of memory bandwidth. By reducing the physical distance data must travel between compute and storage, the design targets the energy and latency costs of data movement that constrain conventional architectures where memory and processors sit on separate dies.
ANALYSIS The 520 TFLOPS figure achieved on a mature 14nm node points to how much performance headroom architectural innovation can unlock independent of lithography advances. The debut arrives while access to advanced foreign chips remains constrained: Beijing recently permitted ByteDance and Tencent to each receive Nvidia H200 chips, the first significant inflow after a months-long lockout[1]. A domestically produced accelerator that sidesteps process-node restrictions could reduce dependence on such limited import windows.