{"dataset":{"id":"56833","dataset_id":"on004408","name":"EEG responses to continuous naturalistic speech\n","description":"This dataset comprises 128-channel EEG recordings from healthy, neurotypical adults listening to naturalistic speech segments from an audiobook version of 'The Old Man and the Sea'. The data were originally collected across two studies investigating cortical entrainment to speech and semantic processing during continuous narrative listening. Accompanying audio stimuli and forced-aligned phoneme/word timing annotations are included to support neurolinguistic analyses.","owner_user_id":15,"status":"active","github_repo":"nemarDatasets/on004408","concept_doi":"10.82901/nemar.on004408","latest_version_doi":"10.82901/nemar.on004408.v1.0.0","created_at":"2026-06-24 07:01:50","updated_at":"2026-08-19 14:38:06","zenodo_concept_id":null,"is_sandbox":0,"visibility":"public","ezid_status":"public","enrichment_json":"{\n  \"version\": \"2.0\",\n  \"pipeline_stage\": \"enriched\",\n  \"title\": \"EEG responses to continuous naturalistic speech\\n\",\n  \"description\": \"This dataset comprises 128-channel EEG recordings from healthy, neurotypical adults listening to naturalistic speech segments from an audiobook version of 'The Old Man and the Sea'. The data were originally collected across two studies investigating cortical entrainment to speech and semantic processing during continuous narrative listening. Accompanying audio stimuli and forced-aligned phoneme/word timing annotations are included to support neurolinguistic analyses.\",\n  \"methods_description\": \"EEG was recorded using a 128-channel BioSemi ActiveTwo system at a sampling rate of 512 Hz while participants listened to segments of an audiobook. Recordings are unfiltered and unreferenced, with each subject's EEG segment start aligned to the corresponding audio stimulus. Word and phoneme timing annotations were generated using the Prosodylab-Aligner forced-alignment software and manually inspected.\",\n  \"license\": \"CC0\",\n  \"dataset_type\": \"raw\",\n  \"authors\": {\n    \"Giovanni M Di Liberto\": {},\n    \"Michael P Broderick\": {},\n    \"Ole Bialas\": {},\n    \"Edmund C Lalor\": {}\n  },\n  \"keywords\": [\n    {\n      \"term\": \"EEG\"\n    },\n    {\n      \"term\": \"naturalistic speech\"\n    },\n    {\n      \"term\": \"Speech Perception\",\n      \"subject_scheme\": \"MeSH\",\n      \"value_uri\": \"http://id.nlm.nih.gov/mesh/D013067\"\n    },\n    {\n      \"term\": \"cortical entrainment\"\n    },\n    {\n      \"term\": \"auditory comprehension\"\n    },\n    {\n      \"term\": \"forced alignment\"\n    }\n  ],\n  \"related_identifiers\": [\n    {\n      \"identifier\": \"https://github.com/nemarDatasets/on004408\",\n      \"identifier_type\": \"URL\",\n      \"relation_type\": \"IsDescribedBy\"\n    },\n    {\n      \"identifier\": \"10.18112/openneuro.ds004408.v1.0.8\",\n      \"identifier_type\": \"DOI\",\n      \"relation_type\": \"IsDerivedFrom\"\n    },\n    {\n      \"identifier\": \"https://nemar.org/dataset/on004408\",\n      \"identifier_type\": \"URL\",\n      \"relation_type\": \"IsDescribedBy\"\n    }\n  ],\n  \"resource_type_general\": \"Dataset\",\n  \"resource_type_specific\": \"EEG Dataset\",\n  \"modalities\": [\n    \"eeg\"\n  ],\n  \"sizes\": [\n    \"20.1 GB (1182 files)\"\n  ],\n  \"formats\": [\n    \".TextGrid\",\n    \".eeg\",\n    \".json\",\n    \".md\",\n    \".tsv\",\n    \".txt\",\n    \".vhdr\",\n    \".vmrk\",\n    \".wav\",\n    \".yml\"\n  ],\n  \"source_hash\": \"194e54350211d38592a7405f30d270a837976dad1d8cb98a05c653a886625ae3\"\n}","last_activity_at":"2026-06-24 07:01:50","source":"openneuro","source_id":"ds004408","subject_count":19,"modalities":"eeg","age_min":null,"age_max":null,"file_size":20083253026,"total_files":1950,"tasks":"listening","metadata_columns_error":null,"staleness_warn_stage":null,"staleness_admin_notified_at":null,"authors":"Giovanni M Di Liberto, Michael P Broderick, Ole Bialas, Edmund C Lalor","license":"CC0","readme":"[![DOI](https://img.shields.io/badge/DOI-10.82901%2Fnemar.on004408-blue)](https://doi.org/10.82901/nemar.on004408)\n\nThe data in one study [^1] and then added to by another [^2] and contains EEG responses of healthy, neurotypical adults who listened to naturalistic speech. The subjects listened to segments from an audio book version of \"The Old Man and the Sea\" and their brain activity was recorded using a 128-channel ActiveTwo EEG system (BioSemi). \n\nThe stimuli folder contains .wav files of the presented audiobook segments as well as a .TextGrid file for each segment, containng the timing of  words and phonemes in that segment. The text grids were generated using the forced-alignment software Prosodylab-Aligner [^3] and inspected by eye. Each subject's folder contains one EEG-recording per audio segment and their starts are aligned (the EEG recordings are longer than the audio to a varying extent).  The recordings are unfiltered, unreferenced and sampled at 512 Hz.\n\n[^1]: Di Liberto, G. M., O’sullivan, J. A., & Lalor, E. C. (2015). Low-frequency cortical entrainment to speech reflects phoneme-level processing. Current Biology, 25(19), 2457-2465.\n\n[^2]: Broderick, M. P., Anderson, A. J., Di Liberto, G. M., Crosse, M. J., & Lalor, E. C. (2018). Electrophysiological correlates of semantic dissimilarity reflect the comprehension of natural, narrative speech. Current Biology, 28(5), 803-809.\n\n[^3]: Gorman, K., Howell, J., & Wagner, M. (2011). Prosodylab-aligner: A tool for forced alignment of laboratory speech. Canadian Acoustics, 39(3), 192-193.\n","bids_version":"1.7.0","sessions_count":null,"publish_date":null,"embedding_dirty":0,"license_tier":"public","zarr_status":"ready","zarr_converted_at":"2026-08-06 22:02:05","zarr_store_count":380,"zarr_index_etag":"3d460dc9a2504cb6ba0c135dc9c860d6","zarr_source_commit":"abb1ee185c5363dd23f50ae1705b32099489a1be","archive_status":"ready","archive_size":16943390900,"archive_retry_count":0,"records_status":"ready","archive_skip_reason":null,"zarr_errors":0,"zarr_failure_count":0,"zarr_deterministic":0,"zarr_failed_at":null,"num_dataset_citations":0,"num_datapaper_citations":90,"n_channels":128,"electrode_system":"biosemi","has_hed":0,"hed_version":null,"is_exemplar":0,"bytes_present":20080811444,"data_complete":1,"withdrawn_at":null,"withdrawn_reason":null,"archive_complete":null,"archive_absent_files":null,"archive_declared_files":null,"zarr_pool_breaks":null,"total_recording_duration":74141.05599999991,"recording_duration_min":168.824,"recording_duration_max":477.668,"recording_count":380,"recordings_unavailable":0,"recordings_measured":380,"channel_count_min":128,"channel_count_max":128,"sampling_frequency":512,"power_line_frequency":null,"eeg_reference":null,"placement_scheme":null,"sweep_stamps":"{\"enrichment_updated_at\":\"2026-08-19 14:37:49\",\"metadata_updated_at\":\"2026-08-19 14:38:04\",\"archive_checked_at\":\"2026-06-24 07:24:36\",\"zarr_checked_at\":null,\"records_checked_at\":\"2026-06-24 07:12:04\",\"citations_updated_at\":\"2026-09-08 03:00:46\",\"channel_montage_checked_at\":\"2026-06-28 23:24:59\",\"hed_checked_at\":\"2026-06-30 04:58:33\",\"data_checked_at\":null,\"availability_report_at\":\"2026-07-23 01:18:44\",\"recording_stats_at\":\"2026-09-02 11:32:45\",\"signal_defaults_at\":\"2026-09-02 12:12:34\"}","participants":19,"num_citations":90,"latest_version":"v1.0.0","zarr_verify_status":null,"zarr_verified_at":null,"owner_username":"nemarAdmin","owner_github":"nemarAdmin","file_size_formatted":"18.70 GB","zarr_data_failures":null,"zarr_index_url":"https://zarr.nemar.org/on004408/zarr/index.json","attestation_deposit_type":null,"attestation_key_status":null,"attestation_deidentified":null,"attestation_no_duplicate":null,"attestation_upstream_source":null,"attestation_accepted_at":null}}