On July 15, 2026, the international standard ITU-T F.746.25:2026 "System Framework and Functional Requirements for In-Vehicle Multimodal Voice Interaction" led by Changan Automobile was officially released on the ITU (International Telecommunication Union) website. As an authoritative agency under the United Nations responsible for communications, ITU is one of the three most authoritative international standards organizations globally, and its specifications have global applicability.

In the field of intelligent cockpits, traditional single-mode in-vehicle voice interaction technology commonly suffers from issues such as inaccurate recognition, false wake-up, and response delays. To solve this problem, Changan Automobile took the lead in initiating a project in 2023, and after three years of refinement and nine rounds of proposals, relying on extensive mass-production cockpit practices, it innovatively launched the "audio-video" multimodal interaction function and established the international standard "System Framework and Functional Requirements for In-Vehicle Multimodal Voice Interaction", filling the global standard gap in the field of in-vehicle multimodal voice interaction. This standard establishes a globally unified layered architecture for in-vehicle interaction, integrating voice, vision, and text multidimensional features, comprehensively optimizing the entire process including voice enhancement, intelligent wake-up, and semantic recognition, thoroughly breaking through the technical bottlenecks of traditional single-mode voice interaction. Compared with traditional voice interaction, Changan's self-developed "audio-video" multimodal interaction function uses multi-channel fusion of listening (voice) + seeing (camera) + sensing (sensors/gestures), supported by multimodal wake-up-free, enhanced command recognition, voice-body fusion understanding, multimodal signal enhancement, and individual control, creating a seamless operation experience that covers all cockpit scenarios such as travel, comfort, entertainment, and security, effectively maintaining the safety bottom line of in-vehicle interaction.

Currently, brands such as Deepal and Changan Qiyuan have fully adopted the "audio-video" multimodal interaction function, with cumulative installations exceeding one million vehicles, covering markets in China, Thailand, Malaysia, Indonesia, etc., effectively solving pain points such as false wake-up, low wake-up rate, response delays, and wake-word dependence in high-noise environments; at the same time, integrating multimodal data of speaker voice and posture gives the voice assistant the ability to "hear" and "see", making in-vehicle dialogue safer, smoother, and more human-like.

This standard release is an important milestone in Changan Automobile's commitment to R&D of core technologies and deep cultivation in the field of intelligent in-vehicle interaction. Changan Automobile adheres to safety as the foundation, starting from the user experience of drivers and passengers, using mature technology refined by massive data to build a globally unified specification, and using warm technology to safeguard every global journey, allowing users worldwide to enjoy a better voice interaction experience.