Huawei Technologies on September 28 open-sourced the pretraining, supervised fine-tuning and reinforcement-learning code behind its openPangu-2.0 model family, giving outside developers the tooling to reproduce and extend the full training pipeline on Ascend hardware1,2.
The release, attributed to Huawei's Ascend Tribe community, comprises two projects: openPangu-2.0-Training, covering pretraining and SFT, and openPangu-2.0-RL, handling post-training reinforcement learning. Both are designed for Huawei's Ascend-based training ecosystem. Prior releases had already placed inference weights for openPangu-2.0-Pro and openPangu-2.0-Flash on Hugging Face Inc.; this round shifts the focus from downloading finished parameters to running the code that produced them.
openPangu-2.0-Pro carries roughly 505 billion total parameters with about 18 billion active per token, a 512K context window and a training corpus of approximately 34 trillion tokens. openPangu-2.0-Flash is smaller, at about 92 billion total parameters with roughly 6 billion active per token, the same 512K context window and a comparable token budget. Both are mixture-of-experts architectures.
Model cards on Hugging Face Inc. describe openPangu as Huawei Technologies' open AI model brand for Ascend-native training and inference.
Architecture notes shared across Pro and Flash list multi-head latent attention, a DSA-plus-SWA layered mix at roughly a 1:2 ratio, a four-branch mHC residual topology, three-head multi-token prediction and Muon-optimizer training. Post-training documentation on both model cards references combined fast/slow SFT, multi-specialized RL and online distillation, abbreviated OPD.
ANALYSIS The code release converts openPangu-2.0 from a weight-only open model into a more fully reproducible artifact, a distinction that matters for teams building on Ascend rather than Nvidia GPUs. By publishing the pretrain-through-RL pipeline, Huawei Technologies is positioning Ascend as a platform others can train on, not just run inference against.