AI News (2026/1/14): Zhipu and Huawei Open-Source the First Domestic Chip-Trained Multimodal SOTA Model GLM-Image
Executive Summary:
Zhipu AI and Huawei have jointly open-sourced the next-generation image generation model GLM-Image, which is the first SOTA multimodal model to be fully trained on the domestic Ascend Atlas 800T A2 chip. The model uses an innovative "autoregressive + diffusion decoder" hybrid architecture and has achieved the best open-source model performance on the CVTG-2K and LongText-Bench benchmarks, particularly excelling in Chinese character generation tasks.
News Details
Zhipu AI and Huawei have jointly open-sourced the next-generation image generation model GLM-Image. This model is the first SOTA multimodal model to be fully trained on the domestic Ascend Atlas 800T A2 chip. GLM-Image adopts an innovative "autoregressive + diffusion decoder" hybrid architecture, which has enabled it to achieve the best open-source model performance on the CVTG-2K and LongText-Bench benchmarks, particularly in Chinese character generation tasks.
Key Points
Autoregressive + Diffusion Decoder: GLM-Image uses an innovative hybrid architecture that combines the strengths of autoregressive models and diffusion decoders. Autoregressive models can generate high-quality local details, while diffusion decoders can handle global structures and complex backgrounds, making the model perform exceptionally well in generating complex images.
Domestic Ascend Atlas 800T A2 Chip: This is the first SOTA multimodal model to be fully trained on the domestic Ascend Atlas 800T A2 chip. The Ascend Atlas 800T A2 chip is characterized by high performance and low power consumption, providing robust computational support for large-scale training.
CVTG-2K and LongText-Bench Benchmarks: GLM-Image has achieved the best open-source model performance on the CVTG-2K and LongText-Bench benchmarks. In particular, GLM-Image can generate high-quality, high-resolution Chinese character images, meeting the needs of various application scenarios.
AI-ALL In-Depth Commentary
The joint open-sourcing of the GLM-Image model by Zhipu AI and Huawei not only demonstrates the strength of domestic chips in high-performance computing but also provides a new solution for multimodal generation tasks. The innovative architecture and superior performance of the model are expected to drive the development of image generation technology, especially in specific areas such as Chinese character generation. Additionally, the successful open-sourcing of GLM-Image provides valuable references and tools for other researchers and developers, contributing to the prosperity and development of the AI ecosystem.
