Difference between revisions of "Datasets"

From rosp
m
m
Line 78: Line 78:
 
|{{n/s}}
 
|{{n/s}}
 
|{{n/s}}
 
|{{n/s}}
|read/colloquial
+
|read, colloquial
 
|2
 
|2
 
|conversation
 
|conversation
Line 146: Line 146:
 
|{{yes}}
 
|{{yes}}
 
|-
 
|-
!SPINE1/SPINE2
+
!SPINE1, SPINE2
 
|2000-2001
 
|2000-2001
 
|military
 
|military
Line 159: Line 159:
 
|US English
 
|US English
 
|1k
 
|1k
|command/colloquial
+
|command, colloquial
 
|1 or 2
 
|1 or 2
 
|{{no}}
 
|{{no}}
Line 180: Line 180:
 
|3 (+1 GSM)  (distant)
 
|3 (+1 GSM)  (distant)
 
|{{no}}
 
|{{no}}
|5 x 200 (Academics) / 5 x 1,000 (Companies)
+
|1000 €
 
|[http://catalog.elra.info/index.php?cPath=37_40 purchase] [http://aurora.hsnr.de/aurora-3/reports.html papers]
 
|[http://catalog.elra.info/index.php?cPath=37_40 purchase] [http://aurora.hsnr.de/aurora-3/reports.html papers]
 
|{{dunno}}
 
|{{dunno}}
Line 186: Line 186:
 
|Finnish, German, Spanish, Danish, Italian
 
|Finnish, German, Spanish, Danish, Italian
 
|{{dunno}}
 
|{{dunno}}
|command (read/digits/keywords/spontaneous)
+
|digits, command, read, spontaneous
 
|1
 
|1
 
|{{no}}
 
|{{no}}
Line 229: Line 229:
 
!RWCP Real Environment Speech and Acoustic Database
 
!RWCP Real Environment Speech and Acoustic Database
 
|2001
 
|2001
|domestic/office
+
|domestic, office
 
|{{dunno}}
 
|{{dunno}}
 
|16000-48000
 
|16000-48000
Line 243: Line 243:
 
|1
 
|1
 
|{{no}}
 
|{{no}}
|real rir/reverb
+
|real rir, reverb
 
|loudspeaker
 
|loudspeaker
 
|various
 
|various
|{{no}}/pivoting arm
+
|no, pivoting arm
 
|stationary background noise
 
|stationary background noise
 
|original
 
|original
Line 264: Line 264:
 
|[http://catalog.elra.info/search.php purchase] [http://www.lrec-conf.org/proceedings/lrec2000/html/summary/373.htm paper]
 
|[http://catalog.elra.info/search.php purchase] [http://www.lrec-conf.org/proceedings/lrec2000/html/summary/373.htm paper]
 
|{{dunno}}
 
|{{dunno}}
|300/language
+
|300 per lang
 
|Multiple
 
|Multiple
 
|{{dunno}}
 
|{{dunno}}
|command (read/digits/keywords/spontaneous)
+
|digits, command, read, spontaneous
 
|1
 
|1
 
|{{no}}
 
|{{no}}
Line 375: Line 375:
 
|US English
 
|US English
 
|12k
 
|12k
|command/digits/read/dialogue
+
|digits, command, read, dialogue
 
|1
 
|1
 
|{{no}}
 
|{{no}}
Line 427: Line 427:
 
|29 h
 
|29 h
 
|86
 
|86
|US/non-native English
+
|US English, non-native English
 
|1k
 
|1k
 
|read
 
|read
Line 526: Line 526:
 
!CHIL Meetings
 
!CHIL Meetings
 
|2004-2007
 
|2004-2007
|seminar/meeting
+
|seminar, meeting
 
|60 h
 
|60 h
 
|44100
 
|44100
Line 537: Line 537:
 
|non-native English
 
|non-native English
 
|{{dunno}}
 
|{{dunno}}
|lecture/meeting
+
|seminar, meeting
 
|3 to 20
 
|3 to 20
|seminar/meeting
+
|seminar, meeting
 
|reverb
 
|reverb
 
|human
 
|human
Line 553: Line 553:
 
!SPEECON
 
!SPEECON
 
|2004-2011
 
|2004-2011
|public space/domestic/office/car
+
|public space, domestic, office, car
 
|{{dunno}}
 
|{{dunno}}
 
|16000
 
|16000
Line 561: Line 561:
 
|[http://catalog.elra.info/search.php purchase] [mailto:diskra@appen.com email] [http://www.lrec-conf.org/proceedings/lrec2002/sumarios/177.htm paper]
 
|[http://catalog.elra.info/search.php purchase] [mailto:diskra@appen.com email] [http://www.lrec-conf.org/proceedings/lrec2002/sumarios/177.htm paper]
 
|{{dunno}}
 
|{{dunno}}
|600/language
+
|600 per lang
 
|Multiple
 
|Multiple
 
|{{dunno}}
 
|{{dunno}}
|command/read/spontaneous
+
|command, read, spontaneous
 
|1
 
|1
 
|{{no}}
 
|{{no}}
Line 634: Line 634:
 
!Aurora-5
 
!Aurora-5
 
|2006
 
|2006
|public spaces/domestic/office/car
+
|public spaces, domestic, office, car
 
|{{dunno}}
 
|{{dunno}}
 
|8000
 
|8000
Line 648: Line 648:
 
|1
 
|1
 
|{{no}}
 
|{{no}}
|real rir/simu/no + simulated telephone channel
+
|no, simulated, real rir
 
|loudspeaker
 
|loudspeaker
 
|{{n/s}}
 
|{{n/s}}
Line 753: Line 753:
 
|US English
 
|US English
 
|2.4k (but transcription is incomplete)
 
|2.4k (but transcription is incomplete)
|command/conversation
+
|command, dialogue
 
|1 to 2
 
|1 to 2
 
|conversation
 
|conversation
Line 761: Line 761:
 
|head
 
|head
 
|car
 
|car
|headset (but problem w/ recording quality)
+
|headset (low quality)
 
|{{no}}
 
|{{no}}
 
|{{yes}} (partial)
 
|{{yes}} (partial)
Line 767: Line 767:
 
|{{no}}
 
|{{no}}
 
|-
 
|-
!SASSEC/SiSEC underdetermined
+
!SASSEC, SiSEC underdetermined
 
|2007-2011
 
|2007-2011
 
|cocktail party
 
|cocktail party
Line 783: Line 783:
 
|3 or 4
 
|3 or 4
 
|full
 
|full
|reverb/real rir/simu
+
|simulated, real rir, reverb
|loudspeaker/{{no}}
+
|no, loudspeaker
 
|fixed
 
|fixed
 
|{{no}}
 
|{{no}}
Line 794: Line 794:
 
|{{no}}
 
|{{no}}
 
|-
 
|-
!MC-WSJ-AV/PASCAL SSC2/2012_MMA/REVERB RealData
+
!MC-WSJ-AV, PASCAL SSC2, 2012_MMA, REVERB RealData
 
|2007-2014
 
|2007-2014
 
|cocktail party
 
|cocktail party
Line 823: Line 823:
 
!CENSREC-4 (Simulated)
 
!CENSREC-4 (Simulated)
 
|2008
 
|2008
|public spaces/domestic/office/car
+
|public spaces, domestic, office, car
 
|{{dunno}}
 
|{{dunno}}
 
|16000
 
|16000
Line 850: Line 850:
 
!CENSREC-4 (Real)
 
!CENSREC-4 (Real)
 
|2008
 
|2008
|public spaces/domestic/office/car
+
|public spaces, domestic, office, car
 
|{{dunno}}
 
|{{dunno}}
 
|16000
 
|16000
Line 940: Line 940:
 
|11 h
 
|11 h
 
|91
 
|91
|US/non-native English
+
|US English, non-native English
 
|5k
 
|5k
 
|colloquial
 
|colloquial
Line 972: Line 972:
 
|1 or 3
 
|1 or 3
 
|full
 
|full
|reverb (other room)/{{no}}
+
|no, reverb (other room)
 
|loudspeaker
 
|loudspeaker
 
|various
 
|various
Line 1,010: Line 1,010:
 
|{{no}}
 
|{{no}}
 
|-
 
|-
!CHiME 1/CHiME 2 Grid
+
!CHiME 1, CHiME 2 Grid
 
|2011-2012
 
|2011-2012
 
|domestic
 
|domestic
Line 1,147: Line 1,147:
 
!REVERB SimData
 
!REVERB SimData
 
|2013
 
|2013
|domestic/office
+
|domestic, office
 
|25 h
 
|25 h
 
|16000
 
|16000

Revision as of 15:58, 8 August 2014

This page aims to provide a list of datasets with detailed attributes and links to corresponding research results (papers, numerical results, output transcriptions, intermediary data, etc). Each dataset may be used for one or more applications: automatic speech recognition, speaker identification and verification, source localization, speech enhancement and separation...

Disclaimer: Only publicly available datasets with a total duration longer than 5 min are listed.

Datasets General attributes Speech Channel Noise Ground truth
release scenario total duration sampling rate degraded channels cameras cost links speech duration unique speakers language unique words speaking style simultaneous speakers speaker overlap channel type radiation speaker location speaker movements noise type speech signal speaker location and orientation words nonverbal traits noise events
ShATR 1994 meeting 37 min 48000 3 (distant) no free download email paper 37 min 5 UK English 1k colloquial 5 multiple conversations reverb human quasi-fixed head meeting headset yes yes no yes
LLSEC 1996 conversation 1.4 h 16000 4 (distant) no free download email ? 12 N/S N/S read, colloquial 2 conversation reverb human quasi-fixed head hallway, restaurant no yes no no no
RWCP Spoken Dialog Corpus 1996-1997 conversation 10 h 16000 2 (close but cross-talk) no free download email paper 10 h 39 Japanese ? colloquial 1 or 2 conversation reverb human quasi-fixed head stationary background noise no no yes no no
Aurora-2 2000 public spaces 33 h 8000-16000 1 (close) no TIDigits download email paper 33 h 214 US English 11 digits 1 no no (simulated telephone channel) human N/S no various real environments original N/S yes no yes
SPINE1, SPINE2 2000-2001 military 38 h 16000 2 (close) no 2 x ($800 (audio) + $500 (transcripts)) + 3 x ($1000 (audio) + $600 (transcripts)) purchase email paper ? 100 US English 1k command, colloquial 1 or 2 no no (simulated transmission channels) human quasi-fixed head military (pre-recorded noise played in sound booth while recording speech) no no yes no no
Aurora-3 (subset of SpeechDat-Car) 2000-2003 car ? 16000 3 (+1 GSM) (distant) no 1000 € purchase papers ? ? Finnish, German, Spanish, Danish, Italian ? digits, command, read, spontaneous 1 no reverb human quasi-fixed head car close-talk no yes no no
RWCP Meeting Speech Corpus 2001 meeting 3.5 h 16000-48000 1 (distant) 3 free download email paper 3.5 h ? Japanese ? colloquial 1 to 5 meeting low reverb human quasi-fixed head stationary background noise headset no yes no no
RWCP Real Environment Speech and Acoustic Database 2001 domestic, office ? 16000-48000 30 (distant) no free download email paper ? 5 Japanese ? read 1 no real rir, reverb loudspeaker various no, pivoting arm stationary background noise original yes yes no yes
SpeechDat-Car 2001-2011 car ? 16000 3 (+1 GSM) (distant) no 1.1 Million for all 10 languages. Each costs 39k to 182k purchase paper ? 300 per lang Multiple ? digits, command, read, spontaneous 1 no reverb human quasi-fixed head car close-talk no yes no no
Aurora-4 2002 public spaces ? 8000-16000 1 (close) no WSJ0 download email paper ? 101 US English 10k read 1 no no (simulated telephone channel) human N/S no various real environments original N/S yes no yes
TED 2002 seminar 47 h 16000 1 (distant) no $275 (audio) + $250 (transcripts) purchase paper 47 h 188 English (mostly non-native) ? lecture 1 or more seminar reverb human quasi-fixed head stationary background noise lapel no yes (partial) no no
CUAVE 2002 cocktail party 3 h 44100 1 (distant) 1 free download email paper 3 h 36 US English 10 digits 1 or 2 full reverb human quasi-fixed head stationary background noise no no yes no no
CU-Move ("Microphone Array Data"; downsampled data with more speakers but less channels exist) 2002-2011 car 286 h 44100 6 to 8 (distant) no $25k with UT-Drive purchase email paper 286 h 172 US English 12k digits, command, read, dialogue 1 no reverb human quasi-fixed head car no no yes no no
CENSREC-1 (Aurora-2J) 2003 public spaces ? 8000 1 (close) no free download email paper 214 Japanese 11 digits 1 no various microphones and simulated channels human N/S no various real environments original N/S yes no yes
AVICAR 2004 car 29 h 16000 7 (distant) 4 free download email paper 29 h 86 US English, non-native English 1k read 1 no reverb human quasi-fixed head car no no yes no no
AV16.3 2004 meeting 1.5 h 16000 16 (distant) 3 free download email paper 1.5 h 12 N/S N/S colloquial 1 to 3 full reverb human various walk stationary background noise no yes no no no
ICSI Meeting Corpus 2004 meeting 72 h 16000 6 (distant) no $1900 (audio) + $900 (transcripts) purchase email paper 72 h 53 US English 13k meeting 3 to 10 meeting reverb human quasi-fixed head stationary background noise headset (some lapel) no yes yes no
NIST Meeting Pilot Corpus Speech 2004 meeting 15 h 16000 7 (distant) no (released but not currently available for download) $4000 (audio) + $1500 (transcripts) purchase email paper 15 h 61 US English 6k meeting 3 to 9 meeting reverb human various walk stationary background noise headset+lapel no yes no no
CHIL Meetings 2004-2007 seminar, meeting 60 h 44100 79 to 147 (distant) 6 to 9 3 500 purchase email paper ? ? non-native English ? seminar, meeting 3 to 20 seminar, meeting reverb human quasi-fixed head meeting (scenarized) headset yes yes yes no
SPEECON 2004-2011 public space, domestic, office, car ? 16000 3 (distant) no 29 x 75000 for all languages purchase email paper ? 600 per lang Multiple ? command, read, spontaneous 1 no reverb human quasi-fixed head various real environments headset no yes no no
CENSREC-2 2005 car ? 16000 1 (distant) no free download email paper ? 214 Japanese 11 digits 1 no reverb human quasi-fixed head car headset no yes no no
CENSREC-3 2005 car ? 16000 1 (distant) no free except phonetically balanced training set: JPY 21000 (Universities) / JPY 105000 (Companies) purchase email paper ? 18 (+293 in training) Japanese 50 in evaluation; unknown but larger in phonetically-balanced utterances of training set read 1 no reverb human quasi-fixed head car headset no yes no no
Aurora-5 2006 public spaces, domestic, office, car ? 8000 1 (distant) no TIDigits download email paper ? 225 US English 11 digits 1 no no, simulated, real rir loudspeaker N/S no various real environments original no yes no yes
AMI 2006 meeting 100 h 16000 16 (distant) 6 free download email paper ? 189 UK English 8k meeting 4 (18% overlap) meeting reverb human quasi-fixed head stationary background noise headset+lapel yes yes yes no
PASCAL SSC 2006 cocktail party 18.5 min (+ 8.5h clean training data) 25000 1 (mixing console) no free email paper 18.5 min (+ 8.5h clean training data) 34 UK English 51 command 2 full no human N/S no no original N/S yes no no
HIWIRE 2007 airplane 21 h 16000 1 (close) no 50 purchase email paper 21 h 81 non-native English 133 command 1 no no human N/S head airplane original N/S yes no no
UT-Drive 2007 car 40 h 25000 5 (distant) 2 $25k with CU-Move download email paper 40 h 25 (more exist but not included in latest release 3.0) US English 2.4k (but transcription is incomplete) command, dialogue 1 to 2 conversation reverb human quasi-fixed head car headset (low quality) no yes (partial) no no
SASSEC, SiSEC underdetermined 2007-2011 cocktail party 19 min 16000 2 (distant) no free download email paper 19 min 16 N/S N/S read 3 or 4 full simulated, real rir, reverb no, loudspeaker fixed no no original+spatial image yes no no no
MC-WSJ-AV, PASCAL SSC2, 2012_MMA, REVERB RealData 2007-2014 cocktail party 10 h 16000 8 to 40 (distant) no $1 500 purchase email paper paper ? 45 UK English 10k read 1 or 2 full reverb human various walk stationary background noise headset+lapel yes yes no no
CENSREC-4 (Simulated) 2008 public spaces, domestic, office, car ? 16000 1 (distant) no free download email paper ? 214 Japanese 11 digits 1 no real rir mouth simulator fixed no various real environments original no yes no yes
CENSREC-4 (Real) 2008 public spaces, domestic, office, car ? 16000 1 (distant) no free download email paper ? 10 Japanese 11 digits 1 no reverb human quasi-fixed head various real environments headset no yes no yes
DICIT 2008 domestic 6 h 48000 16 (distant) 2 free download email paper 1 h ? Italian ? command 4 no reverb human various walk domestic (scenarized) headset+tv yes yes no yes
SiSEC head-geometry 2008 cocktail party 1.9 h 16000 2 (distant) no free download email paper 1.9 h ? N/S N/S read 2 full real rir loudspeaker various no no original+spatial image yes no no no
COSINE 2009 conversation 38 h 48000 20 (distant) no free download email paper 11 h 91 US English, non-native English 5k colloquial 2 to 7 conversation reverb human various walk various real environments headset+throat mic no yes no no
SiSEC real-world noise 2010 public spaces 20 min 16000 2 to 4 (distant) no free download email paper 20 min 6 N/S N/S read 1 or 3 full no, reverb (other room) loudspeaker various no various real environments original+spatial image yes no no no
SiSEC dynamic 2010-2011 cocktail party 11 min 16000 2 to 4 (distant) no free download email paper 11 min ? N/S N/S read Many but only 2 simultaneous simu reverb loudspeaker various simu no original+spatial image yes no no no
CHiME 1, CHiME 2 Grid 2011-2012 domestic 70 h with some overlap 16000 2 (distant) no free download email paper 12 h 34 UK English 51 command 1 no real rir dummy quasi-fixed simu domestic yes yes yes no no
CHiME 2 WSJ0 2012 domestic 78 h with some overlap 16000 2 (distant) no WSJ0 download email paper 33 h 101 US English 11k read 1 no real rir dummy fixed no domestic yes yes yes no no
ETAPE 2012 debates, outdoor interviews, and other TV/radio broadcasts selected for large speaker overlap and/or noise 42 h 16000 1 (mixing console) 1 ? email paper 32 h 347 French 16k colloquial 1 or more (7% overlap on average, up to 10% in debates) conversation some reverb human quasi-fixed head various real environments no N/S yes no yes
GALE (Chinese broadcast conversation) 2013 conversation (TV Broadcast) 120 h 16000 1 (mixing console) no $2000 (audio) + $1500 (transcripts) purchase email 108 h ? Mandarin ? colloquial 1 or more conversation no human quasi-fixed head no no N/S yes no no
GALE (Arabic broadcast conversation) 2013 conversation (TV Broadcast) 251 h 16000 1 (mixing console) no 2 x [$2000 (audio) + $1500 (transcripts)] purchase email 234 h ? Arabic ? colloquial 1 or more conversation no human quasi-fixed head no no N/S yes no no
REVERB SimData 2013 domestic, office 25 h 16000 8 (distant) no WSJCAM0 purchase email paper 25 h 130 UK English 10k read 1 no real rir loudspeaker fixed no experimental room original+spatial image yes yes no yes
DIRHA 2014 domestic 3.8 h 48000 40 (distant) no free download email paper 1.3 h 30 Italian, German, Greek, Portuguese various various 1 or more simu real rir loudspeaker various no domestic (sum of individual noises) yes yes yes no yes

Automatic speech recognition

1st CHiME Challenge (2011)

Artificially distorted version of the small vocabulary GRID audio-visual corpus (audio only). Binaural reverberated speech with speaker situated in front of the microphones. Additive household noises impinging from different directions. Clean-training, noisy-training, development and evaluation sets available, see

Jon Barker, E. Vincent, N. Ma, H. Christensen, P. Green, "The PASCAL CHiME speech separation and recognition challenge", Computer Speech & Language, Volume 27, Issue 3, May 2013, Pages 621-633.

Available from Computer Speech and Language here

Corpus available here (no cost)

Resources

  • Training recipe of the challenge for HTK here.

Baselines

  • See the paper above for results for a wide range of techniques.


AURORA 5 (2007)

Artificially distorted version of the digits TI-DIGITS corpus. Additive noise and additive noise plus reverberant speech sets. Variable SNR range. Various mixed training sets, no evaluation set, see

G. Hirsch "Aurora-5 Experimental Framework for the Performance Evaluation of Speech Recognition in Case of a Hands-free Speech Input in Noisy Environments", Niederrhein University of Applied Sciences, 2007.

Paper available online here (no cost)

Corpus available from LDC here

Resources

  • Training recipe for HTK is provided with the corpora.

Baselines

  • Reproducible baseline: The above cited paper includes a baseline for the ETSI Advanced Front-End.


AURORA 4 (2002)

Artificially distorted version of the 5K word Wall Street Journal corpus (WSJ0). Stationary and non-stationary noises added. Second recordings with distant mismatched microphone. Clean-training, mixed-training, noisy training and test sets available. No evaluation set, see

G. Hirsch "Experimental Framework for the Performance Evaluation of Speech Recognition Front-ends on a Large Vocabulary Task", ETSI STQ Aurora DSR Working Group, 2002.

Paper available with the corpus.

Corpora available from ELRA here and here

Resources

  • Training recipe for HTK available here. Note that this recipe is for Wall-Street Journal (WSJ0), which is the clean speech version of AURORA4. Small changes are needed in the feature extraction scripts to account for different file terminations.

Speaker identification and verification

Speech enhancement and separation

Other applications

Contribute a dataset

To contribute a new dataset, please

  • create an account and login
  • go to the wiki page above corresponding to your application; if it does not exist yet, you may create it
  • click on the "Edit" link at the top of the page and add a new section for your dataset (the datasets are ordered by year of collection)
  • click on the "Save page" link at the bottom of the page to save your modifications

Please make sure to provide the following information:

  • name of the dataset and year of collection
  • authors, institution, contact information
  • link to the dataset and to side resources (lexicon, language model, etc)
  • short description (nature of the data, license, etc) and link to a paper/report describing the dataset, if any
  • at least 1 research result obtained for this dataset (see below)

We currently cannot provide storage space for large datasets. Please upload the dataset at a stable URL on the website of your institution or elsewhere and provide its URL only. If this is not possible, please contact the resources sharing working group.

Contribute a research result

To contribute a new research result, please

  • create an account and login
  • go to the wiki page and the section corresponding to the dataset for which this result was obtained
  • click on the "Edit" link on the right of the section header and add a new item for your result
  • click on the "Save page" link at the bottom of the page to save your modifications

Please make sure to provide the following information:

  • authors, paper/report title, means of publication
  • link to the pdf of the paper
  • link to derived data (output transcriptions, intermediary data, etc)
  • Code and instructions to reproduce experiments (if available)

In order to save storage space, please do not upload the paper on this wiki, but link it as much as possible from your institutional archive, from another public archive (e.g., arxiv) or from the publisher website (e.g., ieexplore).

We currently cannot provide storage space for large datasets. Please upload the derived data at a stable URL on the website of your institution or elsewhere and provide its URL only. If this is not possible, please contact the resources sharing working group.