{"dataset":{"id":"403","dataset_id":"nm000180","name":"Brennan2019: EEG during Naturalistic Listening to Alice in Wonderland","description":"This dataset comprises scalp EEG recordings from 45 participants during passive listening to a naturalistic 12.4-minute audiobook excerpt (first chapter of Alice's Adventures in Wonderland). The study investigates how the brain constructs hierarchical syntactic structure and generates rapid linguistic predictions during continuous speech comprehension. Each word event is annotated with linguistic predictors including lexical frequency, acoustic properties, and surprisal estimates from n-gram, recurrent neural network, and context-free grammar models.","owner_user_id":19,"status":"active","github_repo":"nemarDatasets/nm000180","concept_doi":"10.82901/nemar.nm000180","latest_version_doi":"10.82901/nemar.nm000180.v1.1.3","created_at":"2026-04-08 21:14:22","updated_at":"2026-07-10 21:52:16","zenodo_concept_id":"20499947","is_sandbox":0,"visibility":"public","ezid_status":"public","enrichment_json":"{\n  \"version\": \"2.0\",\n  \"pipeline_stage\": \"validated\",\n  \"title\": \"Brennan2019: EEG during Naturalistic Listening to Alice in Wonderland\",\n  \"description\": \"This dataset comprises scalp EEG recordings from 45 participants during passive listening to a naturalistic 12.4-minute audiobook excerpt (first chapter of Alice's Adventures in Wonderland). The study investigates how the brain constructs hierarchical syntactic structure and generates rapid linguistic predictions during continuous speech comprehension. Each word event is annotated with linguistic predictors including lexical frequency, acoustic properties, and surprisal estimates from n-gram, recurrent neural network, and context-free grammar models.\",\n  \"methods_description\": \"EEG was recorded using a 61-channel BrainVision actiCHamp amplifier with an equidistant montage (easyCAP M10), yielding 60 scalp channels plus 1 bipolar vertical EOG and 1 audio channel. Sampling rate was 500 Hz with online filtering (0.01–200 Hz). Participants listened to audiobook segments presented via insert earphones (Etymotic ER-2) at 45 dB above individual hearing threshold. Data were collected between March 2015 and December 2016 at the University of Michigan Computational Neurolinguistics Lab.\",\n  \"license\": \"CC-BY-4.0\",\n  \"dataset_type\": \"raw\",\n  \"authors\": {\n    \"Jonathan R. Brennan\": {\n      \"orcid\": \"0000-0002-3639-350X\"\n    },\n    \"John T. Hale\": {}\n  },\n  \"keywords\": [\n    {\n      \"term\": \"EEG\"\n    },\n    {\n      \"term\": \"naturalistic listening\"\n    },\n    {\n      \"term\": \"language comprehension\"\n    },\n    {\n      \"term\": \"syntactic structure\"\n    },\n    {\n      \"term\": \"linguistic prediction\"\n    },\n    {\n      \"term\": \"surprisal\"\n    },\n    {\n      \"term\": \"auditory processing\"\n    }\n  ],\n  \"related_identifiers\": [\n    {\n      \"identifier\": \"https://github.com/nemarDatasets/nm000180\",\n      \"identifier_type\": \"URL\",\n      \"relation_type\": \"IsDescribedBy\"\n    },\n    {\n      \"identifier\": \"10.21105/joss.01896\",\n      \"identifier_type\": \"DOI\",\n      \"relation_type\": \"References\"\n    },\n    {\n      \"identifier\": \"10.1038/s41597-019-0104-8\",\n      \"identifier_type\": \"DOI\",\n      \"relation_type\": \"References\"\n    },\n    {\n      \"identifier\": \"10.1371/journal.pone.0207741\",\n      \"identifier_type\": \"DOI\",\n      \"relation_type\": \"IsDerivedFrom\"\n    },\n    {\n      \"identifier\": \"10.7302/746w-g237\",\n      \"identifier_type\": \"DOI\",\n      \"relation_type\": \"IsDerivedFrom\"\n    },\n    {\n      \"identifier\": \"10.7302/Z29C6VNH\",\n      \"identifier_type\": \"DOI\",\n      \"relation_type\": \"IsDerivedFrom\"\n    },\n    {\n      \"identifier\": \"10.1098/rstb.2019.0305\",\n      \"identifier_type\": \"DOI\",\n      \"relation_type\": \"References\"\n    },\n    {\n      \"identifier\": \"https://nemar.org/dataset/nm000180\",\n      \"identifier_type\": \"URL\",\n      \"relation_type\": \"IsDescribedBy\"\n    }\n  ],\n  \"resource_type_general\": \"Dataset\",\n  \"resource_type_specific\": \"EEG Dataset\",\n  \"modalities\": [\n    \"eeg\"\n  ],\n  \"sizes\": [\n    \"4.1 GB (91 files)\"\n  ],\n  \"formats\": [\n    \".csv\",\n    \".doc\",\n    \".eeg\",\n    \".json\",\n    \".mat\",\n    \".md\",\n    \".py\",\n    \".sfp\",\n    \".tsv\",\n    \".txt\",\n    \".vhdr\",\n    \".vmrk\",\n    \".wav\",\n    \".yml\"\n  ],\n  \"source_hash\": \"0646520eb4dd86bc0214f855ac22d00ee8113dcfa6b552433789042bf43e4e37\"\n}","last_activity_at":"2026-04-08 21:14:38","source":null,"source_id":null,"subject_count":45,"modalities":"eeg","age_min":null,"age_max":null,"file_size":4086497797,"total_files":91,"tasks":"alicelistening","metadata_columns_error":null,"staleness_warn_stage":null,"staleness_admin_notified_at":null,"authors":"Jonathan R. Brennan, John T. Hale","license":"CC-BY-4.0","readme":"[![DOI](https://img.shields.io/badge/DOI-10.82901%2Fnemar.nm000180-blue)](https://doi.org/10.82901/nemar.nm000180)\n\nBrennan2019: EEG during Naturalistic Listening to Alice in Wonderland\n====================================================================\n\nOverview\n--------\nScalp EEG recorded while participants listened to a 12.4-minute audiobook\nrecording of the first chapter of *Alice's Adventures in Wonderland* (Lewis\nCarroll). This is a naturalistic auditory language-comprehension paradigm\ndesigned to study how the brain builds **hierarchical syntactic structure** and\nmakes **rapid, incremental linguistic predictions** during continuous listening.\n\nThis BIDS dataset contains **45 participants** (`sub-001` … `sub-048`). The\noriginal study by Brennan & Hale (2019) recorded 49 participants (S01–S49) and\nanalyzed 33 after quality-control exclusions; the mapping from `sub-XXX` to the\noriginal `SXX` identifiers is preserved in the per-subject filenames.\n\nTask\n----\n- **Task label:** `alicelistening`\n- Passive listening to continuous naturalistic speech (an audiobook), presented\n  over insert earphones (Etymotic ER-2) at 45 dB above each participant's\n  individually determined hearing threshold.\n- The ~12.4-minute audio was divided into **12 segments**; a digital trigger was\n  sent at the onset of each segment. (Note: segment triggers are absent from the\n  re-hosted continuous `.vmrk`; word onsets in `events.tsv` are reconstructed\n  from the original per-subject preprocessing output — see *Events*.)\n- The audiobook segments are provided in **`stimuli/`**\n  (`DownTheRabbitHoleFinal_SoundFile1–12.wav`, 16-bit mono 44.1 kHz), and each word\n  event references its segment via the `stim_file` column of `events.tsv`.\n- After listening, participants answered multiple-choice comprehension\n  questions (8 for most subjects; 4 for an early subset — see the\n  `comprehension_n_total` column). Per-subject scores are in `participants.tsv`\n  (`sourcedata/comprehension-scores.txt`).\n\nRecording Setup\n---------------\n- **Amplifier:** BrainVision actiCHamp (Brain Products GmbH)\n- **Electrodes:** 61-channel equidistant montage (easyCAP M10), of which **60\n  scalp EEG channels** are retained here, plus **1 bipolar vertical EOG (VEOG)**\n  over the left eye and **1 audio channel (AUD)** carrying a digitized copy of\n  the acoustic stimulus.\n- **Sampling rate:** 500 Hz\n- **Online filter:** 0.01–200 Hz\n- **Reference:** average reference (offline)\n- Data collected March 2015 – December 2016 at the University of Michigan\n  Computational Neurolinguistics Lab (CNL Lab).\n\nEvents and Linguistic Annotations\n---------------------------------\nThe scientific value of this dataset is the alignment of the continuous EEG to\nthe **word-by-word linguistic structure** of the narrative. Per-subject\n`*_events.tsv` files provide one event per spoken word (2129 words), with onsets\nin the subject's own EEG time base, and are annotated with the predictors used\nin Brennan & Hale (2019):\n\n| Column | Description |\n|---|---|\n| `onset`, `duration` | Word onset/duration in the EEG recording (seconds) |\n| `trial_type` | `word` |\n| `word` | The spoken word (orthographic form) |\n| `word_index` | Order of the word in the story (Brennan `Order`; 1–2150, 2129 annotated with 21 gaps) |\n| `sentence_id` | Sentence number the word belongs to |\n| `position_in_sentence` | Ordinal position of the word within its sentence |\n| `segment` | Audio segment (1–12) the word occurs in |\n| `is_lexical` | 1 = open-class/content word, 0 = closed-class/function word |\n| `log_freq` | Log lexical frequency of the word |\n| `log_freq_prev`, `log_freq_next` | Log frequency of the preceding / following word |\n| `sound_power` | Acoustic sound power at word onset |\n| `ngram_surprisal` | Surprisal from a 3-gram language model (`NGRAM`) |\n| `rnn_surprisal` | Surprisal from a recurrent neural-network language model (`RNN`) |\n| `cfg_surprisal` | Surprisal from a context-free-grammar parser (`CFG`) |\n\nThe `ngram_surprisal`, `rnn_surprisal`, and `cfg_surprisal` predictors index\nprogressively more hierarchical linguistic structure and are the core measures\ntested in the paper. Annotations derive from `AliceChapterOne-EEG.csv` (word list\n+ predictors) joined to the per-subject `proc.mat` word-onset samples.\n\nSource Data\n-----------\n`sourcedata/` re-hosts the upstream annotation and preprocessing files from the\nDeep Blue Data record (v2, DOI `10.7302/746w-g237`):\n- `AliceChapterOne-EEG.csv` — 2129-word linguistic annotations\n- `proc/` — per-subject preprocessing outputs (`SXX.mat`; word-onset `trl`)\n- `easycapM10-acti61_elec.sfp` — electrode template positions\n- `comprehension-questions.doc`, `comprehension-scores.txt` — behavioral task\n- `datasets.mat` — subject index\n\nHow to cite\n-----------\n> Brennan, J. R., & Hale, J. T. (2019). Hierarchical structure guides rapid\n> linguistic predictions during naturalistic listening. *PLoS ONE*, 14(1),\n> e0207741. https://doi.org/10.1371/journal.pone.0207741\n\n> Brennan, J. R. (2023). EEG Datasets for Naturalistic Listening to \"Alice in\n> Wonderland\" (v2). University of Michigan – Deep Blue Data.\n> https://doi.org/10.7302/746w-g237\n\nRelated\n-------\n> Bhattasali, S., Brennan, J., Luh, W.-M., Franzluebbers, B., & Hale, J. (2020).\n> The Alice datasets: fMRI & EEG observations of natural language comprehension.\n> *LREC 2020*. https://aclanthology.org/2020.lrec-1.15/\n\n> Brennan, J. R., & Martin, A. E. (2020). Phase synchronization varies\n> systematically with linguistic structure composition. *Phil. Trans. R. Soc. B*,\n> 375. https://doi.org/10.1098/rstb.2019.0305\n\nBIDS references\n---------------\nAppelhoff, S., et al. (2019). MNE-BIDS. *JOSS* 4:1896. https://doi.org/10.21105/joss.01896\n\nPernet, C. R., et al. (2019). EEG-BIDS. *Scientific Data* 6:103. https://doi.org/10.1038/s41597-019-0104-8\n","bids_version":"1.9.0","sessions_count":0,"publish_date":"2026-06-01 23:19:53","embedding_dirty":0,"license_tier":"attribution","zarr_status":"ready","zarr_converted_at":"2026-07-25 01:44:09","zarr_store_count":45,"zarr_index_etag":"d398f71c975501b77cdd4ba882855fe4","zarr_source_commit":"03a81ca87bdbabb044dbc22ec95b596275432898","archive_status":"ready","archive_size":3388392259,"archive_retry_count":0,"records_status":null,"archive_skip_reason":null,"zarr_errors":0,"zarr_failure_count":0,"zarr_deterministic":0,"zarr_failed_at":null,"num_dataset_citations":0,"num_datapaper_citations":127,"n_channels":60,"electrode_system":"other","has_hed":0,"hed_version":null,"is_exemplar":0,"bytes_present":null,"data_complete":null,"withdrawn_at":null,"withdrawn_reason":null,"archive_complete":null,"archive_absent_files":null,"archive_declared_files":null,"zarr_pool_breaks":null,"total_recording_duration":32954.988,"recording_duration_min":726.648,"recording_duration_max":746.752,"recording_count":45,"recordings_unavailable":0,"recordings_measured":45,"channel_count_min":62,"channel_count_max":62,"sampling_frequency":500,"power_line_frequency":60,"eeg_reference":"average (offline)","placement_scheme":"easycap-M10","sweep_stamps":"{\"enrichment_updated_at\":\"2026-07-03 10:40:06\",\"metadata_updated_at\":\"2026-07-03 10:40:06\",\"archive_checked_at\":\"2026-06-05 01:32:59\",\"zarr_checked_at\":\"2026-06-07 17:58:26\",\"records_checked_at\":null,\"citations_updated_at\":\"2026-09-03 03:00:56\",\"channel_montage_checked_at\":\"2026-06-28 22:55:54\",\"hed_checked_at\":\"2026-06-30 04:15:54\",\"data_checked_at\":null,\"availability_report_at\":\"2026-07-23 01:07:01\",\"recording_stats_at\":\"2026-09-02 11:31:47\",\"signal_defaults_at\":\"2026-09-02 11:41:58\"}","participants":45,"num_citations":127,"latest_version":"v1.0.0","zarr_verify_status":null,"zarr_verified_at":null,"owner_username":"bruaristimunha","owner_github":"bruAristimunha","file_size_formatted":"3.81 GB","zarr_data_failures":null,"zarr_index_url":"https://zarr.nemar.org/nm000180/zarr/index.json","attestation_deposit_type":null,"attestation_key_status":null,"attestation_deidentified":null,"attestation_no_duplicate":null,"attestation_upstream_source":null,"attestation_accepted_at":null}}