{"dataset":{"id":"421","dataset_id":"nm000253","name":"Wang et al. 2024 — Brain Treebank: Large-scale intracranial recordings from naturalistic language stimuli","description":"Brain Treebank is a large-scale intracranial EEG dataset comprising 43 hours of iEEG recordings from 10 epilepsy patients watching naturalistic Hollywood movies, with 1,688 electrodes sampled at 2048 Hz. The dataset includes time-aligned linguistic annotations with word-level transcripts and Universal Dependencies syntax trees, providing a unique resource for studying neural language processing during naturalistic stimulation.","owner_user_id":19,"status":"active","github_repo":"nemarDatasets/nm000253","concept_doi":"10.82901/nemar.nm000253","latest_version_doi":"10.82901/nemar.nm000253.v1.0.0","created_at":"2026-04-12 19:16:23","updated_at":"2026-07-10 22:01:43","zenodo_concept_id":"20587480","is_sandbox":0,"visibility":"public","ezid_status":"public","enrichment_json":"{\n  \"version\": \"2.0\",\n  \"pipeline_stage\": \"validated\",\n  \"authors\": {\n    \"Christopher Wang\": {},\n    \"Adam Yaari\": {},\n    \"Aaditya K Singh\": {},\n    \"Vighnesh Subramaniam\": {},\n    \"Dana Rosenfarb\": {},\n    \"Jan DeWitt\": {},\n    \"Pranav Misra\": {},\n    \"Joseph R Madsen\": {},\n    \"Scellig Stone\": {},\n    \"Gabriel Kreiman\": {},\n    \"Boris Katz\": {},\n    \"Ignacio Cases\": {},\n    \"Andrei Barbu\": {}\n  },\n  \"related_identifiers\": [\n    {\n      \"identifier\": \"10.48550/arXiv.2411.08343\",\n      \"identifier_type\": \"DOI\",\n      \"relation_type\": \"IsDerivedFrom\"\n    },\n    {\n      \"identifier\": \"https://github.com/nemarDatasets/nm000253\",\n      \"identifier_type\": \"URL\",\n      \"relation_type\": \"IsDescribedBy\"\n    },\n    {\n      \"identifier\": \"https://nemar.org/dataexplorer/detail?dataset_id=nm000253\",\n      \"identifier_type\": \"URL\",\n      \"relation_type\": \"IsDescribedBy\"\n    }\n  ],\n  \"title\": \"Wang et al. 2024 — Brain Treebank: Large-scale intracranial recordings from naturalistic language stimuli\",\n  \"license\": \"CC BY 4.0\",\n  \"dataset_type\": \"raw\",\n  \"resource_type_general\": \"Dataset\",\n  \"modalities\": [\n    \"ieeg\"\n  ],\n  \"resource_type_specific\": \"iEEG Dataset\",\n  \"sizes\": [\n    \"545.9 GB (81 files)\"\n  ],\n  \"formats\": [\n    \".eeg\",\n    \".json\",\n    \".md\",\n    \".py\",\n    \".sh\",\n    \".tsv\",\n    \".vhdr\",\n    \".vmrk\",\n    \".yml\",\n    \".zip\"\n  ],\n  \"description\": \"Brain Treebank is a large-scale intracranial EEG dataset comprising 43 hours of iEEG recordings from 10 epilepsy patients watching naturalistic Hollywood movies, with 1,688 electrodes sampled at 2048 Hz. The dataset includes time-aligned linguistic annotations with word-level transcripts and Universal Dependencies syntax trees, providing a unique resource for studying neural language processing during naturalistic stimulation.\",\n  \"methods_description\": \"Intracranial electrophysiological recordings were acquired at 2048 Hz from an average of 168 electrodes per subject during passive viewing of Hollywood movies. Audio transcripts were manually annotated for word onsets and parsed into Universal Dependencies syntax trees. Data were converted to BIDS-iEEG format (BrainVision, IEEE_FLOAT_32, multiplexed) with electrode coordinates and per-subject electrode information provided.\",\n  \"keywords\": [\n    {\n      \"term\": \"intracranial EEG\"\n    },\n    {\n      \"term\": \"stereo-EEG\"\n    },\n    {\n      \"term\": \"language processing\"\n    },\n    {\n      \"term\": \"naturalistic stimuli\"\n    },\n    {\n      \"term\": \"epilepsy\"\n    },\n    {\n      \"term\": \"electrocorticography\"\n    },\n    {\n      \"term\": \"syntax\"\n    },\n    {\n      \"term\": \"linguistic annotations\"\n    },\n    {\n      \"term\": \"Universal Dependencies\"\n    },\n    {\n      \"term\": \"naturalistic language\"\n    }\n  ],\n  \"source_hash\": \"c78e6e4485458ee284b03849f09a9465a1d1dd57c0e122a96f6475126a2cd4d3\"\n}","last_activity_at":"2026-06-04 07:40:11","source":null,"source_id":null,"subject_count":10,"modalities":"ieeg","age_min":null,"age_max":null,"file_size":545856861968,"total_files":81,"tasks":"movie","metadata_columns_error":null,"staleness_warn_stage":null,"staleness_admin_notified_at":null,"authors":"Christopher Wang, Adam Yaari, Aaditya K Singh, Vighnesh Subramaniam, Dana Rosenfarb, Jan DeWitt, Pranav Misra, Joseph R Madsen, Scellig Stone, Gabriel Kreiman, Boris Katz, Ignacio Cases, Andrei Barbu","license":"CC BY 4.0","readme":"[![DOI](https://img.shields.io/badge/DOI-10.82901%2Fnemar.nm000253-blue)](https://doi.org/10.82901/nemar.nm000253)\n\n# Brain Treebank: large-scale intracranial (iEEG) recordings from naturalistic language stimuli\n\n## Summary\n\nThe Brain Treebank is a large-scale dataset of intracranial electrophysiological (iEEG /\nstereo-EEG) recordings collected from **10 epilepsy patients** (at Boston Children's\nHospital) while they watched Hollywood movies. Recordings were acquired at **2048 Hz**\nfrom on average **168 electrodes per subject (1,688 electrodes total)**, totalling\nroughly **43 hours across 26 trials** (one trial = one movie). The audio of each movie\nwas transcribed and word onsets manually annotated, and each transcript was parsed into\nUniversal Dependencies (UD) syntax trees — making this one of the largest datasets of\nintracranial recordings grounded in naturalistic language.\n\nThis NEMAR record provides the dataset converted to **BIDS-iEEG** (BrainVision) format:\neach trial is one BIDS run (`task-movie`, `run-01` …), with signals in `sub-*/ieeg/`.\n\n## Modality and paradigm\n\n- **Modality:** Intracranial EEG / stereo-EEG (iEEG-BIDS), 2048 Hz, BrainVision\n  (IEEE_FLOAT_32, multiplexed)\n- **Task / paradigm:** Passive naturalistic viewing of Hollywood movies (`task-movie`),\n  with time-aligned word-level language annotations (transcripts and UD syntax trees in\n  `code/transcripts.zip` and `code/trees.zip`)\n- **Population:** 10 epilepsy patients undergoing intracranial monitoring\n\n## Participants and data structure\n\n- **10 subjects** (`sub-01` … `sub-10`), 26 movie-viewing runs in total.\n- Electrode labels and per-subject electrode information are in each `sub-*/ieeg/`\n  (`electrodes.tsv`, `coordsystem.json`); corrupted electrodes are flagged as `bad` in\n  `channels.tsv` (from the Brain Treebank `corrupted_elec.json`).\n- Electrode coordinates are provided as-is from the Brain Treebank `localization.zip`\n  (`code/localization.zip`); see each `coordsystem.json` for important caveats about\n  units and the absence of a published transform to a standard template.\n\n## Original dataset / data paper\n\nPlease cite the original publication when using this dataset:\n\n> Wang, C., Yaari, A. U., Singh, A. K., Subramaniam, V., Rosenfarb, D., DeWitt, J.,\n> Misra, P., Madsen, J. R., Stone, S., Kreiman, G., Katz, B., Cases, I., & Barbu, A.\n> (2024). *Brain Treebank: Large-scale intracranial recordings from naturalistic language\n> stimuli.* Advances in Neural Information Processing Systems 37 (NeurIPS 2024, Datasets\n> and Benchmarks Track). arXiv:2411.08343.\n\n- **Preprint / DOI:** [arXiv:2411.08343](https://doi.org/10.48550/arXiv.2411.08343)\n- **NeurIPS 2024 proceedings:** https://proceedings.neurips.cc/paper_files/paper/2024/hash/aefa2385b3f33abf1526ae4e2c208cd9-Abstract-Datasets_and_Benchmarks_Track.html\n- **Project site / source data:** https://braintreebank.dev/\n- **Original code release:** https://github.com/czlwang/brain_treebank_code_release\n- **Ethics:** Boston Children's Hospital / Harvard IRB; all subjects gave informed consent.\n\n## BIDS conversion\n\nPer-trial HDF5 recordings from braintreebank.dev were converted to BIDS-iEEG (BrainVision)\nwith the EEGDash conversion script in `code/convert_braintreebank.py`. Each trial (one\nmovie) becomes one BIDS run. EEG-BIDS / MNE-BIDS were used only for standardisation; the\ndata themselves are from the original Brain Treebank release. Please credit the original\ncreators (Wang et al.) and cite the paper above.\n\n## License\n\nCC BY 4.0 (see `dataset_description.json`).\n","bids_version":"1.9.0","sessions_count":0,"publish_date":"2026-04-12 19:16:23","embedding_dirty":0,"license_tier":"attribution","zarr_status":"ready","zarr_converted_at":"2026-09-07 13:32:18","zarr_store_count":2,"zarr_index_etag":"fe4b1552433bc9e38b0883baf2cfda43","zarr_source_commit":"7fe9b07131937b08ab07c211ec9ed08948571a84","archive_status":null,"archive_size":null,"archive_retry_count":0,"records_status":null,"archive_skip_reason":"Archive (>100 GB) removed to save space; use the per-file direct download (#752).","zarr_errors":24,"zarr_failure_count":0,"zarr_deterministic":0,"zarr_failed_at":"2026-09-07 13:32:18","num_dataset_citations":0,"num_datapaper_citations":0,"n_channels":null,"electrode_system":null,"has_hed":0,"hed_version":"8.2.0","is_exemplar":0,"bytes_present":null,"data_complete":null,"withdrawn_at":null,"withdrawn_reason":null,"archive_complete":null,"archive_absent_files":null,"archive_declared_files":null,"zarr_pool_breaks":0,"total_recording_duration":14047.398000000001,"recording_duration_min":5646.023,"recording_duration_max":8401.375,"recording_count":26,"recordings_unavailable":24,"recordings_measured":2,"channel_count_min":218,"channel_count_max":218,"sampling_frequency":null,"power_line_frequency":null,"eeg_reference":null,"placement_scheme":null,"sweep_stamps":"{\"enrichment_updated_at\":\"2026-06-08 03:33:02\",\"metadata_updated_at\":\"2026-06-08 03:33:02\",\"archive_checked_at\":\"2026-06-17 01:02:40\",\"zarr_checked_at\":null,\"records_checked_at\":null,\"citations_updated_at\":\"2026-09-08 03:00:53\",\"channel_montage_checked_at\":\"2026-06-28 23:00:45\",\"hed_checked_at\":\"2026-06-30 04:30:51\",\"data_checked_at\":null,\"availability_report_at\":\"2026-07-23 01:09:36\",\"signal_defaults_at\":\"2026-09-02 11:49:03\",\"recording_stats_at\":\"2026-09-08 03:00:48\"}","participants":10,"num_citations":0,"latest_version":"v1.0.0","zarr_verify_status":null,"zarr_verified_at":null,"owner_username":"bruaristimunha","owner_github":"bruAristimunha","file_size_formatted":"508 GB","zarr_data_failures":{"count":0,"detail_ref":"zarr/index.json","pending":24,"discovered":26},"zarr_index_url":"https://zarr.nemar.org/nm000253/zarr/index.json","attestation_deposit_type":null,"attestation_key_status":null,"attestation_deidentified":null,"attestation_no_duplicate":null,"attestation_upstream_source":null,"attestation_accepted_at":null}}