DeepZang, China's first large language model for the Tibetan language, is advancing Tibetan language processing and helping enhance the intelligent development of ethnic languages, according to stakeholders at a recent conference on the model's latest upgrade in Hohhot, the capital of the Inner Mongolia autonomous region.
The conference focused on advances in Tibetan artificial intelligence technology and multilingual simultaneous interpretation, with the aim of promoting the intelligent development of ethnic languages across the country.
Developed by local company Choknor, DeepZang was launched in March and facilitates communication in Tibetan, standard Chinese and English, offering functions such as AI conversations, instant translation and speech-to-text conversion.
The model has so far amassed nearly 70 million entries of standardized parallel corpora and more than 30,500 hours of voice data covering Tibetan's three major dialects — Utsang, Kham and Amdo — creating what its developer describes as China's largest and most accurately annotated Tibetan speech database to date.
Tenzin Norbu, founder of Choknor, said ethnic languages such as Tibetan and Mongolian face common challenges due to limited resources, and called for cross-language and cross-regional technological collaboration to advance the informatization of ethnic languages in China.
"Adapting general large models to the Tibetan language so that people of all ethnic groups can equally benefit from digital technology development has been a research focus for both industry and academia for many years," Tenzin Norbu said.
Siqin Tu, a PhD candidate specializing in AI engineering at Inner Mongolia Normal University, said Tibetan is transitioning from digital preservation to digital-intelligent reconstruction through AI technology.
The World Record Certification Agency previously recognized DeepZang as "The World's First Tibetan Large Language Model", acknowledging its pioneering status globally.
After several months of operation, DeepZang has made significant progress in dialect adaptation and semantic understanding, marking a crucial step from being "usable" to becoming more "user-friendly", Tenzin Norbu said.
The company is working to expand the application of Tibetan intelligent technologies to public services, cultural heritage, education, medical services and cultural tourism in ethnic regions, he said.
Since its deployment, DeepZang has been applied in a range of fields.
Lobsang Donyo, a Tibetan-Chinese bilingual translator in Lhokha, Xizang autonomous region, said AI-assisted translation has reduced the time needed to complete a document from 40 minutes with a three-person team to just over 20 minutes for one person.
Academic users have also highlighted its potential. Sonam Yontan, a PhD candidate at Xizang University, said the model has improved research efficiency.
"Its translation and search features are practical and handy," he said. "We can sort through documents and locate references far more quickly."
He added that the model represents an unprecedented breakthrough for Tibetan in the AI field.
In healthcare, DeepZang helps address language barriers between "aid-Xizang" doctors and local Tibetan herders, providing a more efficient solution for communication.
It has also been used in cultural heritage-related work, including drafting yak-trading contracts and composing Tibetan poetry while maintaining cultural authenticity.
"The model has attracted nearly 400,000 users and recorded more than 350 million visits, with token usage exceeding 23 billion," Tenzin Norbu said. "Users aged between 18 and 40 account for more than 70 percent of its audience, with a significant user base in Xizang and the provinces of Qinghai, Sichuan and Gansu, as well as other regions."
The conference also emphasized that the true value of large language models lies in serving the public and improving people's livelihoods.
