MAGICDATA Mandarin Chinese Read Speech Corpus
Summary: The corpus by Magic Data Technology Co., Ltd. , containing 755 hours of scripted read speech data from 1080 native speakers of the Mandarin Chinese spoken in mainland China. The sentence transcription accuracy is higher than 98%.
License: Attribution-NonCommercial-NoDerivatives 4.0 International Public License (CC BY-NC-ND 4.0)
Downloads (use a mirror closer to you):
train_set.tar.gz [52G] ( Training set speech and transcripts ) Mirrors: [US] [EU] [CN]
dev_set.tar.gz [1.0G] (Development set speech and transcripts ) Mirrors: [US] [EU] [CN]
test_set.tar.gz [2.2G] (Test set speech and transcripts ) Mirrors: [US] [EU] [CN]
metadata.tar.gz [3.8M] (supplementary resources, incl. data introduction (in English and Chinese) and speaker information ) Mirrors: [US] [EU] [CN]
About this resource:
The contents and the corresponding descriptions of the corpus include:
- The corpus contains 755 hours of speech data, which is mostly mobile recorded data.
- 1080 speakers from different accent areas in China are invited to participate in the recording.
- The sentence transcription accuracy is higher than 98%.
- Recordings are conducted in a quiet indoor environment.
- The database is divided into training set, validation set, and testing set in a ratio of 51: 1: 2.
- Detail information such as speech data coding and speaker information is preserved in the metadata file.
- The domain of recording texts is diversified, including interactive Q&A, music search, SNS messages, home command and control, etc.
- Segmented transcripts are also provided.
The corpus is a subset of a much bigger data ( 10566.9 hours Chinese Mandarin Speech Corpus ) set which was recorded in the same environment. Please feel free to contact us via firstname.lastname@example.org for more details.
Please cite the corpus as "Magic Data Technology Co., Ltd., "http://www.imagicdatatech.com/index.php/home/dataopensource/data_info/id/101", 05/2019".
Magic Data Technology Co., Ltd. (referred to as Magic Data) was established in 2016. Through our higher-expertise and higher-precision data services, Magic Data has quickly grown into one of the foremost companies in artificial intelligence industry. We strive to provide the most efficient and highest quality one-stop data services for customers in the fields of speech recognition, intelligent imaging and Natural Language Understanding (NLU). Our services include data scheme design, data collection, data annotation/transcription, etc.
- Tel: (+86) 10-82527250
- Email: email@example.com
External URL: http://www.imagicdatatech.com/index.php/home/dataopensource/data_info/id/101 Full description from the company website