{"dataset":{"id":"414","dataset_id":"nm000229","name":"Gwilliams et al. 2023 — Introducing MEG-MASC: a high-quality magneto-encephalography dataset for evaluating natural speech processing","description":"MEG-MASC is a high-quality magnetoencephalography dataset comprising raw MEG recordings from 27 English speakers listening to approximately two hours of naturalistic stories from the Manually Annotated Sub-Corpus (MASC). The dataset includes precise temporal annotations of word and phoneme onsets/offsets, organized according to the Brain Imaging Data Structure (BIDS) standard. This benchmark dataset enables large-scale encoding and decoding analyses of neural responses to natural speech processing, with accompanying code for validation analyses including temporal decoding of phonetic features and word frequency effects.","owner_user_id":19,"status":"active","github_repo":"nemarDatasets/nm000229","concept_doi":"10.82901/nemar.nm000229","latest_version_doi":"10.82901/nemar.nm000229.v1.0.1","created_at":"2026-04-11 11:22:48","updated_at":"2026-07-10 21:58:32","zenodo_concept_id":"20519659","is_sandbox":0,"visibility":"public","ezid_status":"public","enrichment_json":"{\n  \"version\": \"2.0\",\n  \"pipeline_stage\": \"validated\",\n  \"authors\": {\n    \"Laura Gwilliams\": {\n      \"orcid\": \"0000-0002-9213-588X\"\n    },\n    \"Graham Flick\": {},\n    \"Alec Marantz\": {},\n    \"Liina Pylkkänen\": {},\n    \"David Poeppel\": {\n      \"orcid\": \"0000-0003-0184-163X\"\n    },\n    \"Jean-Rémi King\": {}\n  },\n  \"related_identifiers\": [\n    {\n      \"identifier\": \"10.1038/s41597-023-02752-5\",\n      \"identifier_type\": \"DOI\",\n      \"relation_type\": \"IsDerivedFrom\"\n    },\n    {\n      \"identifier\": \"https://github.com/nemarDatasets/nm000229\",\n      \"identifier_type\": \"URL\",\n      \"relation_type\": \"IsDescribedBy\"\n    },\n    {\n      \"identifier\": \"https://nemar.org/dataexplorer/detail?dataset_id=nm000229\",\n      \"identifier_type\": \"URL\",\n      \"relation_type\": \"IsDescribedBy\"\n    }\n  ],\n  \"funding_references\": [\n    {\n      \"funder_name\": \"Dingwall Foundation (L. Gwilliams)\"\n    },\n    {\n      \"funder_name\": \"Agence Nationale de la Recherche\",\n      \"award_number\": \"ANR-17-EURE-0017\"\n    },\n    {\n      \"funder_name\": \"Abu Dhabi Research Institute\",\n      \"award_number\": \"G1001\"\n    },\n    {\n      \"funder_name\": \"Dingwall Foundation\"\n    }\n  ],\n  \"title\": \"Gwilliams et al. 2023 — Introducing MEG-MASC: a high-quality magneto-encephalography dataset for evaluating natural speech processing\",\n  \"license\": \"CC0\",\n  \"dataset_type\": \"raw\",\n  \"resource_type_general\": \"Dataset\",\n  \"modalities\": [\n    \"meg\",\n    \"anat\"\n  ],\n  \"resource_type_specific\": \"MEG Dataset\",\n  \"sizes\": [\n    \"106.6 GB (471 files)\"\n  ],\n  \"formats\": [\n    \".con\",\n    \".gz\",\n    \".json\",\n    \".md\",\n    \".mrk\",\n    \".pos\",\n    \".py\",\n    \".sh\",\n    \".tsv\",\n    \".txt\",\n    \".wav\",\n    \".yml\"\n  ],\n  \"description\": \"MEG-MASC is a high-quality magnetoencephalography dataset comprising raw MEG recordings from 27 English speakers listening to approximately two hours of naturalistic stories from the Manually Annotated Sub-Corpus (MASC). The dataset includes precise temporal annotations of word and phoneme onsets/offsets, organized according to the Brain Imaging Data Structure (BIDS) standard. This benchmark dataset enables large-scale encoding and decoding analyses of neural responses to natural speech processing, with accompanying code for validation analyses including temporal decoding of phonetic features and word frequency effects.\",\n  \"methods_description\": \"Participants listened to four fictional stories presented across two ~1-hour MEG sessions (5 subjects completed 1 session). Stories were synthesized using Mac OS Mojave text-to-speech with varying female voices (n=3) and speech rates (145-205 words per minute). Word and phoneme timing were inferred using forced-alignment between audio and text files via the 'gentle aligner' from the lowerquality Python module. Each recording session included story presentations intermixed with random word lists and comprehension questions.\",\n  \"keywords\": [\n    {\n      \"term\": \"Magnetoencephalography\",\n      \"subject_scheme\": \"MeSH\",\n      \"value_uri\": \"http://id.nlm.nih.gov/mesh/D015225\"\n    },\n    {\n      \"term\": \"MEG\"\n    },\n    {\n      \"term\": \"speech processing\"\n    },\n    {\n      \"term\": \"natural language\"\n    },\n    {\n      \"term\": \"phonetic decoding\"\n    },\n    {\n      \"term\": \"brain imaging\"\n    },\n    {\n      \"term\": \"BIDS\"\n    }\n  ],\n  \"source_hash\": \"4e3dd4d9f8b82325c63b2ea56fc2ccf71a6b2eb60ff403e0df357a3071a75245\"\n}","last_activity_at":"2026-04-13 13:04:11","source":null,"source_id":null,"subject_count":27,"modalities":"anat,meg","age_min":18,"age_max":41,"file_size":106631832692,"total_files":471,"tasks":"0,1,2,3","metadata_columns_error":null,"staleness_warn_stage":null,"staleness_admin_notified_at":null,"authors":"Laura Gwilliams, Graham Flick, Alec Marantz, Liina Pylkkänen, David Poeppel, Jean-Rémi King","license":"CC0","readme":"[![DOI](https://img.shields.io/badge/DOI-10.82901%2Fnemar.nm000229-blue)](https://doi.org/10.82901/nemar.nm000229)\n\n# MEG-MASC: a high-quality magneto-encephalography dataset for evaluating natural speech processing.\n\nLaura Gwilliams, Graham Flick, Alec Marantz, Liina Pylkkänen, David Poeppel, Jean-Rémi King\n\n- [Paper](https://arxiv.org)\n- [Data](https://osf.io/rguwj/)\n- [Code](https://github.com/kingjr/meg-masc)\n\n## Abstract\nThe \"MEG-MASC\" dataset provides a curated set of raw magnetoencephalography (MEG) recordings of 27 English speakers who listened to two hours of naturalistic stories. Each participant performed two identical sessions, involving listening to four fictional stories from the Manually Annotated Sub-Corpus (MASC) intermixed with random word lists and comprehension questions. We time-stamp the onset and offset of each word and phoneme in the metadata of the recording, and organize the dataset according to the 'Brain Imaging Data Structure' (BIDS). This data collection provides a suitable benchmark to large-scale encoding and decoding analyses of temporally-resolved brain responses to speech. We provide the Python code to replicate several validations analyses of the MEG evoked related fields such as the temporal decoding of phonetic features and word frequency. All code and MEG, audio and text data are publicly available to keep with best practices in transparent and reproducible research.\n\n## Please cite\n@article{gwilliams2022neural,\n  title={Neural dynamics of phoneme sequences reveal position-invariant code for content and order},\n  author={Gwilliams, Laura and King, Jean-Remi and Marantz, Alec and Poeppel, David},\n  journal={Nature Communications},\n  volume={13},\n  number={1},\n  pages={1--14},\n  year={2022},\n  publisher={Nature Publishing Group}\n}\n\n## Task organisation\nEach subject listened to four unique stories:\n  - task-0 : 'lw1',\n  - task-1 : 'cable_spool_fort',\n  - task-2 : 'easy_money',\n  - task-3 : 'The_Black_Widow'\n\nStories were presented in a different order to each participant:\n\n  participant_id\t:\ttask_order\n  sub-01\t:\t[0, 1, 2, 3]\n  sub-02\t:\t[0, 1, 3, 2]\n  sub-03\t:\t[0, 2, 3, 1]\n  sub-04\t:\t[3, 0, 1, 2]\n  sub-05\t:\t[2, 3, 1, 0]\n  sub-06\t:\t[0, 2, 1, 3]\n  sub-07\t:\t[0, 3, 1, 2]\n  sub-08\t:\t[3, 1, 0, 2]\n  sub-09\t:\t[2, 1, 3, 0]\n  sub-10\t:\t[1, 2, 3, 0]\n  sub-11\t:\t[1, 3, 2, 0]\n  sub-12\t:\t[2, 0, 3, 1]\n  sub-13\t:\t[1, 3, 0, 2]\n  sub-14\t:\t[1, 0, 3, 2]\n  sub-15\t:\t[2, 1, 0, 3]\n  sub-16\t:\t[3, 0, 2, 1]\n  sub-17\t:\t[1, 2, 3, 0]\n  sub-18\t:\t[2, 0, 1, 3]\n  sub-19\t:\t[0, 3, 2, 1]\n  sub-20\t:\t[2, 3, 0, 1]\n  sub-21\t:\t[1, 2, 3, 0]\n  sub-22\t:\t[1, 0, 2, 3]\n  sub-23\t:\t[0, 2, 3, 1]\n  sub-24\t:\t[3, 1, 2, 0]\n  sub-25\t:\t[0, 1, 3, 2]\n  sub-26\t:\t[3, 1, 0, 2]\n  sub-27\t:\t[1, 2, 3, 0]\n\n\n## Stimulus timestamps\n\nThe timing of each phoneme and each word is provided in each sub-*_ses-*_task-*_events.tsv file, for each subject, session and task. The timing links the MEG recording to the relevant speech moments of that story.\n\nEach events file contains five columns:\n  - onset (float) : onset time of event in seconds\n  - duration (float) : duration of event in seconds\n  - trial_type (dict) : dictionary of key:value pairs providing information about the event\n  - sample (int) : onset time of event in MEG samples\n\n## Stories.\n\nEach participant listened to four fictional stories, over the course of two ~1h-long MEG sessions, with the exception of 5 subjects who only underwent 1 session. The stories were played in different orders across participants. These stories were originally selected because they had been annotated for their syntactic structures (MASC).  The corresponding text files can be found in stimuli/text/*.txt\n\n## Word lists and pseudo-words.\n\nTo potentially investigate MEG responses to words independently of their narrative context, the text of these stories have been supplemented with word lists. Specifically, a random word list consisting of the unique content words (nouns, proper nouns, verbs, adverbs and adjectives) selected from the preceding text segment was added in a random order. In addition, a small fraction (<1%) of non-words were inserted into the natural sentences of the stories. The corresponding text files can be found in stimuli/text_with_wordlist/*.txt. For simplicity, the brain responses to these word lists and to these pseudo words are fully discarded from the present study.\n\n## Audio synthesis.\n\nEach of these stories was synthesized with Mac OS Mojave © version 10.14 text-to-speech. Voices (n=3 female) and speech rates (145 - 205 words per minute) varied every 5-20 sentences. The inter-sentence interval randomly varied between 0 and 1,000 ms. Both speech rate and inter-sentence intervals were sampled from a uniform distribution. Each `text_with_wordlist` files was divided into ~3 min sound files, which can be found in stimuli/audio/*.wav.\n\n## Forced Alignment.\n\nThe timing of words and phonemes were inferred from the forced-alignment between the wav and text files, using the ‘gentle aligner’ from the Python module lowerquality (https://github.com/lowerquality/gentle). We discarded the words that did not get a forced alignment through this procedure. Analysis of the Mel spectrogram and of the phonetic decoding led to better results when using gentle than when using the Penn Forced Aligner originally used in Gwilliams et al MASC. The timing of each word and phoneme can be found in the events.tsv of each individual recording session.\n\n## Verification. To verify that the forced alignment did not have a systematic bias, we systematically check the MEG decoding of phonetic features for each sound file separately.\n\n\n﻿References\n----------\nAppelhoff, S., Sanderson, M., Brooks, T., Vliet, M., Quentin, R., Holdgraf, C., Chaumon, M., Mikulan, E., Tavabi, K., Höchenberger, R., Welke, D., Brunner, C., Rockhill, A., Larson, E., Gramfort, A. and Jas, M. (2019). MNE-BIDS: Organizing electrophysiological data into the BIDS format and facilitating their analysis. Journal of Open Source Software 4: (1896). https://doi.org/10.21105/joss.01896\n\nNiso, G., Gorgolewski, K. J., Bock, E., Brooks, T. L., Flandin, G., Gramfort, A., Henson, R. N., Jas, M., Litvak, V., Moreau, J., Oostenveld, R., Schoffelen, J., Tadel, F., Wexler, J., Baillet, S. (2018). MEG-BIDS, the brain imaging data structure extended to magnetoencephalography. Scientific Data, 5, 180110. https://doi.org/10.1038/sdata.2018.110\n","bids_version":"1.9.0","sessions_count":2,"publish_date":"2026-04-11 11:22:48","embedding_dirty":0,"license_tier":"public","zarr_status":"ready","zarr_converted_at":"2026-09-07 07:42:22","zarr_store_count":159,"zarr_index_etag":"87edc7653202fc8816c7589cd8621311","zarr_source_commit":"0f5a2bd43bd561ef484212b514df74e675dc09ba","archive_status":"ready","archive_size":78499353869,"archive_retry_count":0,"records_status":null,"archive_skip_reason":null,"zarr_errors":37,"zarr_failure_count":0,"zarr_deterministic":0,"zarr_failed_at":"2026-09-07 07:42:22","num_dataset_citations":0,"num_datapaper_citations":46,"n_channels":null,"electrode_system":null,"has_hed":0,"hed_version":"8.2.0","is_exemplar":0,"bytes_present":null,"data_complete":null,"withdrawn_at":null,"withdrawn_reason":null,"archive_complete":null,"archive_absent_files":null,"archive_declared_files":null,"zarr_pool_breaks":1,"total_recording_duration":206192,"recording_duration_min":358,"recording_duration_max":1966,"recording_count":196,"recordings_unavailable":0,"recordings_measured":196,"channel_count_min":257,"channel_count_max":257,"sampling_frequency":null,"power_line_frequency":null,"eeg_reference":null,"placement_scheme":null,"sweep_stamps":"{\"enrichment_updated_at\":\"2026-06-03 03:27:02\",\"metadata_updated_at\":\"2026-06-03 03:27:39\",\"archive_checked_at\":\"2026-06-05 01:33:09\",\"zarr_checked_at\":\"2026-06-07 17:58:33\",\"records_checked_at\":null,\"citations_updated_at\":\"2026-09-08 03:00:47\",\"channel_montage_checked_at\":null,\"hed_checked_at\":\"2026-06-30 04:29:19\",\"data_checked_at\":null,\"availability_report_at\":\"2026-07-23 01:08:47\",\"signal_defaults_at\":\"2026-09-02 11:46:31\"}","participants":27,"num_citations":46,"latest_version":"v1.0.1","zarr_verify_status":null,"zarr_verified_at":null,"owner_username":"bruaristimunha","owner_github":"bruAristimunha","file_size_formatted":"99.31 GB","zarr_data_failures":{"count":0,"detail_ref":"zarr/index.json","pending":37,"discovered":196},"zarr_index_url":"https://zarr.nemar.org/nm000229/zarr/index.json","attestation_deposit_type":null,"attestation_key_status":null,"attestation_deidentified":null,"attestation_no_duplicate":null,"attestation_upstream_source":null,"attestation_accepted_at":null}}