{"dataset":{"id":"61113","dataset_id":"on005574","name":"The \"Podcast\" ECoG dataset","description":"This dataset comprises intracranial electrocorticography (ECoG) recordings from 9 epilepsy patients implanted with grid, depth, and strip electrodes (1,330 electrodes total), collected while participants listened to a 30-minute naturalistic story containing over 5,000 words. It provides raw and minimally preprocessed (high-gamma band) neural data along with aligned auditory stimuli, word-level transcripts, and linguistic features spanning low-level acoustics to large language model embeddings. The dataset is intended to support research on natural language comprehension using high-fidelity invasive recordings and includes tutorials replicating prior findings.","owner_user_id":15,"status":"active","github_repo":"nemarDatasets/on005574","concept_doi":"10.82901/nemar.on005574","latest_version_doi":"10.82901/nemar.on005574.v1.0.0","created_at":"2026-06-27 12:31:00","updated_at":"2026-08-19 01:25:25","zenodo_concept_id":null,"is_sandbox":0,"visibility":"public","ezid_status":"public","enrichment_json":"{\n  \"version\": \"2.0\",\n  \"pipeline_stage\": \"enriched\",\n  \"title\": \"The \\\"Podcast\\\" ECoG dataset\",\n  \"description\": \"This dataset comprises intracranial electrocorticography (ECoG) recordings from 9 epilepsy patients implanted with grid, depth, and strip electrodes (1,330 electrodes total), collected while participants listened to a 30-minute naturalistic story containing over 5,000 words. It provides raw and minimally preprocessed (high-gamma band) neural data along with aligned auditory stimuli, word-level transcripts, and linguistic features spanning low-level acoustics to large language model embeddings. The dataset is intended to support research on natural language comprehension using high-fidelity invasive recordings and includes tutorials replicating prior findings.\",\n  \"methods_description\": \"ECoG data were recorded from 9 participants implanted with grid, depth, and strip electrodes while they listened to a 30-minute naturalistic story. A minimally preprocessed version in the high-gamma spectral band is provided alongside raw data, along with aligned auditory stimuli, word-level transcripts, and linguistic features including acoustic properties and LLM-derived embeddings.\",\n  \"license\": \"CC0\",\n  \"dataset_type\": \"raw\",\n  \"authors\": {\n    \"Zaid Zada\": {},\n    \"Samuel A. Nastase\": {},\n    \"Bobbi Aubrey\": {},\n    \"Itamar Jalon\": {},\n    \"Ariel Goldstein\": {},\n    \"Sebastian Michelmann\": {},\n    \"Haocheng Wang\": {},\n    \"Liat Hasenfratz\": {},\n    \"Werner Doyle\": {},\n    \"Daniel Friedman\": {},\n    \"Patricia Dugan\": {},\n    \"Lucia Melloni\": {},\n    \"Sasha Devore\": {},\n    \"Orrin Devinsky\": {},\n    \"Adeen Flinker\": {},\n    \"Uri Hasson\": {}\n  },\n  \"keywords\": [\n    {\n      \"term\": \"ECoG\"\n    },\n    {\n      \"term\": \"electrocorticography\"\n    },\n    {\n      \"term\": \"language comprehension\"\n    },\n    {\n      \"term\": \"naturalistic stimuli\"\n    },\n    {\n      \"term\": \"intracranial electrophysiology\"\n    },\n    {\n      \"term\": \"high-gamma activity\"\n    },\n    {\n      \"term\": \"large language models\"\n    },\n    {\n      \"term\": \"natural language processing\"\n    },\n    {\n      \"term\": \"narrative comprehension\"\n    },\n    {\n      \"term\": \"story listening\"\n    }\n  ],\n  \"related_identifiers\": [\n    {\n      \"identifier\": \"https://github.com/nemarDatasets/on005574\",\n      \"identifier_type\": \"URL\",\n      \"relation_type\": \"IsDescribedBy\"\n    },\n    {\n      \"identifier\": \"https://nemar.org/dataset/on005574\",\n      \"identifier_type\": \"URL\",\n      \"relation_type\": \"IsDescribedBy\"\n    },\n    {\n      \"identifier\": \"10.18112/openneuro.ds005574.v1.0.2\",\n      \"identifier_type\": \"DOI\",\n      \"relation_type\": \"IsDerivedFrom\"\n    }\n  ],\n  \"funding_references\": [\n    {\n      \"funder_name\": \"National Institutes of Health\",\n      \"award_number\": \"DP1HD091948\"\n    },\n    {\n      \"funder_name\": \"National Institutes of Health\",\n      \"award_number\": \"R01NS109367\"\n    }\n  ],\n  \"resource_type_general\": \"Dataset\",\n  \"resource_type_specific\": \"Structural MRI Dataset\",\n  \"modalities\": [\n    \"anat\",\n    \"ieeg\"\n  ],\n  \"sizes\": [\n    \"15.5 GB (90 files)\"\n  ],\n  \"formats\": [\n    \".csv\",\n    \".edf\",\n    \".fif\",\n    \".gz\",\n    \".hdf5\",\n    \".html\",\n    \".ipynb\",\n    \".json\",\n    \".log\",\n    \".md\",\n    \".py\",\n    \".pyc\",\n    \".sh\",\n    \".tsv\",\n    \".wav\",\n    \".yml\"\n  ],\n  \"source_hash\": \"0f264c3cddd5337478d4df569f0d6fed5dbe186f15dd6e382cecba4d75275ba8\"\n}","last_activity_at":"2026-06-27 12:31:00","source":"openneuro","source_id":"ds005574","subject_count":9,"modalities":"anat,ieeg","age_min":null,"age_max":null,"file_size":15507134455,"total_files":160,"tasks":"podcast","metadata_columns_error":null,"staleness_warn_stage":null,"staleness_admin_notified_at":null,"authors":"Zaid Zada, Samuel A. Nastase, Bobbi Aubrey, Itamar Jalon, Ariel Goldstein, Sebastian Michelmann, Haocheng Wang, Liat Hasenfratz, Werner Doyle, Daniel Friedman, Patricia Dugan, Lucia Melloni, Sasha Devore, Orrin Devinsky, Adeen Flinker, Uri Hasson","license":"CC0","readme":"[![DOI](https://img.shields.io/badge/DOI-10.82901%2Fnemar.on005574-blue)](https://doi.org/10.82901/nemar.on005574)\n\nThe \"Podcast\" ECoG dataset for modeling neural activity during natural story listening.\n\nWe introduce the “Podcast” electrocorticography (ECoG) dataset for modeling neural activity supporting natural narrative comprehension. This dataset combines the exceptional spatiotemporal resolution of human intracranial electrophysiology with a naturalistic experimental paradigm for language comprehension. In addition to the raw data, we provide a minimally preprocessed version in the high-gamma spectral band to showcase a simple pipeline and to make it easier to use. Furthermore, we include the auditory stimuli, an aligned word-level transcript, and linguistic features ranging from low-level acoustic properties to large language model (LLM) embeddings. We also include tutorials replicating previous findings and serve as a pedagogical resource and a springboard for new research. The dataset comprises 9 participants with 1,330 electrodes, including grid, depth, and strip electrodes. The participants listened to a 30-minute story with over 5,000 words. By using a natural story with high-fidelity, invasive neural recordings, this dataset offers a unique opportunity to investigate language comprehension.","bids_version":"1.10.0","sessions_count":null,"publish_date":null,"embedding_dirty":0,"license_tier":"public","zarr_status":"failed","zarr_converted_at":"2026-07-09 11:22:55","zarr_store_count":24,"zarr_index_etag":"3c634fcf57af38bcbfcb1517839723e1","zarr_source_commit":"de611e0f03b3fd8bf3f8a669eb582abbb99dba22","archive_status":"ready","archive_size":13296111984,"archive_retry_count":0,"records_status":"ready","archive_skip_reason":null,"zarr_errors":9,"zarr_failure_count":7,"zarr_deterministic":0,"zarr_failed_at":"2026-08-27 08:21:16","num_dataset_citations":4,"num_datapaper_citations":0,"n_channels":null,"electrode_system":null,"has_hed":0,"hed_version":null,"is_exemplar":0,"bytes_present":15504925732,"data_complete":1,"withdrawn_at":null,"withdrawn_reason":null,"archive_complete":null,"archive_absent_files":null,"archive_declared_files":null,"zarr_pool_breaks":0,"total_recording_duration":null,"recording_duration_min":null,"recording_duration_max":null,"recording_count":null,"recordings_unavailable":null,"recordings_measured":null,"channel_count_min":null,"channel_count_max":null,"sampling_frequency":null,"power_line_frequency":null,"eeg_reference":null,"placement_scheme":null,"sweep_stamps":"{\"enrichment_updated_at\":\"2026-08-19 01:25:16\",\"metadata_updated_at\":\"2026-08-19 01:25:24\",\"archive_checked_at\":\"2026-06-27 12:48:53\",\"zarr_checked_at\":null,\"records_checked_at\":\"2026-06-27 12:40:15\",\"citations_updated_at\":\"2026-09-08 03:00:51\",\"channel_montage_checked_at\":\"2026-06-28 23:49:00\",\"hed_checked_at\":\"2026-06-30 05:22:09\",\"data_checked_at\":\"2026-08-23 03:00:36\",\"availability_report_at\":\"2026-07-23 01:25:48\",\"recording_stats_at\":null,\"signal_defaults_at\":\"2026-09-02 12:38:45\"}","participants":9,"num_citations":4,"latest_version":"v1.0.0","zarr_verify_status":null,"zarr_verified_at":null,"owner_username":"nemarAdmin","owner_github":"nemarAdmin","file_size_formatted":"14.44 GB","zarr_data_failures":{"count":7,"detail_ref":"zarr/index.json","compacted_by":"migration_0074"},"zarr_index_url":null,"attestation_deposit_type":null,"attestation_key_status":null,"attestation_deidentified":null,"attestation_no_duplicate":null,"attestation_upstream_source":null,"attestation_accepted_at":null}}