{"dataset":{"id":"419","dataset_id":"nm000251","name":"He et al. 2025 — VocalMind: A Stereotactic EEG Dataset for Vocalized, Mimed, and Imagined Speech in Tonal Language","description":"VocalMind is a stereotactic EEG (sEEG) dataset comprising intracranial recordings from one participant performing vocalized, mimed, and imagined speech tasks in Mandarin Chinese, a tonal language. The dataset supports research on speech decoding, brain-computer interfaces, and neural mechanisms of overt and covert speech production. Raw recordings sampled at 1000 Hz were converted to BIDS iEEG format with event markers indicating stimulus onsets for each speech modality.","owner_user_id":19,"status":"active","github_repo":"nemarDatasets/nm000251","concept_doi":"10.82901/nemar.nm000251","latest_version_doi":"10.82901/nemar.nm000251.v1.0.0","created_at":"2026-04-12 13:29:00","updated_at":"2026-07-10 22:01:28","zenodo_concept_id":"20738180","is_sandbox":0,"visibility":"public","ezid_status":"public","enrichment_json":"{\n  \"version\": \"2.0\",\n  \"pipeline_stage\": \"validated\",\n  \"title\": \"He et al. 2025 — VocalMind: A Stereotactic EEG Dataset for Vocalized, Mimed, and Imagined Speech in Tonal Language\",\n  \"description\": \"VocalMind is a stereotactic EEG (sEEG) dataset comprising intracranial recordings from one participant performing vocalized, mimed, and imagined speech tasks in Mandarin Chinese, a tonal language. The dataset supports research on speech decoding, brain-computer interfaces, and neural mechanisms of overt and covert speech production. Raw recordings sampled at 1000 Hz were converted to BIDS iEEG format with event markers indicating stimulus onsets for each speech modality.\",\n  \"methods_description\": \"Stereotactic EEG recordings were obtained while the participant performed three speech modes: vocalized, mimed, and imagined speech in Mandarin Chinese. Raw per-trial recordings at 1000 Hz were converted to BIDS iEEG BrainVision format using the EEGDash conversion pipeline, with per-trial files concatenated by speech mode and task type, and events.tsv files marking stimulus onsets.\",\n  \"license\": \"CC BY 4.0\",\n  \"dataset_type\": \"raw\",\n  \"authors\": {\n    \"Tianyu He\": {},\n    \"Mingyi Wei\": {},\n    \"Ruicong Wang\": {},\n    \"Renzhi Wang\": {},\n    \"Shiwei Du\": {},\n    \"Siqi Cai\": {\n      \"orcid\": \"0000-0003-3282-9246\"\n    },\n    \"Wei Tao\": {},\n    \"Haizhou Li\": {}\n  },\n  \"keywords\": [\n    {\n      \"term\": \"stereotactic EEG\"\n    },\n    {\n      \"term\": \"intracranial EEG\"\n    },\n    {\n      \"term\": \"speech decoding\"\n    },\n    {\n      \"term\": \"brain-computer interfaces\"\n    },\n    {\n      \"term\": \"imagined speech\"\n    },\n    {\n      \"term\": \"tonal language\"\n    },\n    {\n      \"term\": \"Mandarin Chinese\"\n    }\n  ],\n  \"related_identifiers\": [\n    {\n      \"identifier\": \"10.5281/zenodo.14696348\",\n      \"identifier_type\": \"DOI\",\n      \"relation_type\": \"IsDerivedFrom\"\n    },\n    {\n      \"identifier\": \"https://github.com/nemarDatasets/nm000251\",\n      \"identifier_type\": \"URL\",\n      \"relation_type\": \"IsDescribedBy\"\n    },\n    {\n      \"identifier\": \"https://nemar.org/dataexplorer/detail?dataset_id=nm000251\",\n      \"identifier_type\": \"URL\",\n      \"relation_type\": \"IsDescribedBy\"\n    },\n    {\n      \"identifier\": \"10.1038/s41597-025-04741-2\",\n      \"identifier_type\": \"DOI\",\n      \"relation_type\": \"IsDerivedFrom\"\n    }\n  ],\n  \"funding_references\": [\n    {\n      \"funder_name\": \"National Natural Science Foundation of China\"\n    },\n    {\n      \"funder_name\": \"Shenzhen Science and Technology Innovation Program\"\n    }\n  ],\n  \"resource_type_general\": \"Dataset\",\n  \"resource_type_specific\": \"iEEG Dataset\",\n  \"modalities\": [\n    \"ieeg\"\n  ],\n  \"sizes\": [\n    \"2.0 GB (332 files)\"\n  ],\n  \"formats\": [\n    \".eeg\",\n    \".json\",\n    \".md\",\n    \".py\",\n    \".tsv\",\n    \".vhdr\",\n    \".vmrk\",\n    \".wav\",\n    \".yml\"\n  ],\n  \"source_hash\": \"4f15be316db3bff45f77b50ea031644d3196079ed4bcc8dca2a2f4480cf5f3cc\"\n}","last_activity_at":"2026-06-03 17:55:32","source":null,"source_id":null,"subject_count":1,"modalities":"ieeg","age_min":null,"age_max":null,"file_size":2030647596,"total_files":332,"tasks":"imagined,mimed,vocalized","metadata_columns_error":null,"staleness_warn_stage":null,"staleness_admin_notified_at":null,"authors":"Tianyu He, Mingyi Wei, Ruicong Wang, Renzhi Wang, Shiwei Du, Siqi Cai, Wei Tao, Haizhou Li","license":"CC BY 4.0","readme":"[![DOI](https://img.shields.io/badge/DOI-10.82901%2Fnemar.nm000251-blue)](https://doi.org/10.82901/nemar.nm000251)\n\n# VocalMind: a stereotactic EEG (sEEG) dataset for vocalized, mimed, and imagined speech in a tonal language\n\n## Summary\n\nVocalMind is a **stereotactic EEG (sEEG / intracranial EEG)** dataset for speech\ndecoding research. Intracranial recordings were obtained while the participant produced\nspeech in a **tonal language (Mandarin Chinese)** under three speech modes — **vocalized**,\n**mimed**, and **imagined** speech — supporting research on speech brain–computer\ninterfaces (BCI) and neural decoding of overt and covert speech.\n\nThe BIDS conversion in this NEMAR record contains intracranial EEG data organised under\nthe `ieeg/` modality, with three tasks corresponding to the speech modes:\n`task-vocalized`, `task-mimed`, and `task-imagined`. It currently includes **1 subject**\n(`sub-01`).\n\n## Modality and paradigm\n\n- **Modality:** Intracranial / stereotactic EEG (sEEG, iEEG-BIDS)\n- **Tasks / paradigms:** Vocalized, mimed, and imagined speech in a tonal language\n  (`task-vocalized`, `task-mimed`, `task-imagined`)\n- **Application:** Speech decoding, speech BCI, tonal-language phonetics\n\n## Participants\n\nThis BIDS dataset includes **1 subject** (`sub-01`). See `participants.tsv` and the data\npaper for clinical/recording details.\n\n## Original dataset / data paper\n\nPlease cite the original Scientific Data descriptor when using this dataset:\n\n> He, T., Wei, M., Wang, R., Wang, R., Du, S., Cai, S., Tao, W., & Li, H. (2025).\n> *VocalMind: A Stereotactic EEG Dataset for Vocalized, Mimed, and Imagined Speech in\n> Tonal Language.* **Scientific Data, 12, 657.**\n> https://doi.org/10.1038/s41597-025-04741-2\n\n- **DOI:** [10.1038/s41597-025-04741-2](https://doi.org/10.1038/s41597-025-04741-2)\n- **Source data:** [Zenodo record 14696348](https://zenodo.org/records/14696348)\n- **Project page:** https://tianyu-h42.github.io/sEEG_speechDecoding/\n\n## Data structure / BIDS conversion\n\nRaw per-trial recordings (1000 Hz) were converted to BIDS iEEG (BrainVision) format with\nthe EEGDash conversion pipeline; per-trial files are concatenated per speech-mode/task\nwith `events.tsv` marking each stimulus onset.\n\n## Attribution\n\nAll data were collected by the original authors (Tianyu He and colleagues). Please credit\nthe original creators and cite the data paper above. This NEMAR record redistributes the\ndataset in BIDS format; MNE-BIDS / iEEG-BIDS and related BIDS tools were used only for\nstandardisation, not as the source of the data.\n\n## License\n\nCC BY 4.0 (see `dataset_description.json`).\n","bids_version":"1.9.0","sessions_count":0,"publish_date":"2026-04-12 13:29:00","embedding_dirty":0,"license_tier":"attribution","zarr_status":"ready","zarr_converted_at":"2026-09-06 02:15:54","zarr_store_count":6,"zarr_index_etag":"a5d7d5dc7132dcd6336db660f4add314","zarr_source_commit":"a82f3a3c042d7ac1af65e863f431d8146afb897e","archive_status":"ready","archive_size":1202214022,"archive_retry_count":0,"records_status":"ready","archive_skip_reason":null,"zarr_errors":0,"zarr_failure_count":0,"zarr_deterministic":0,"zarr_failed_at":null,"num_dataset_citations":0,"num_datapaper_citations":15,"n_channels":null,"electrode_system":null,"has_hed":0,"hed_version":"8.2.0","is_exemplar":0,"bytes_present":null,"data_complete":null,"withdrawn_at":null,"withdrawn_reason":null,"archive_complete":null,"archive_absent_files":null,"archive_declared_files":null,"zarr_pool_breaks":0,"total_recording_duration":4071,"recording_duration_min":357,"recording_duration_max":1000,"recording_count":6,"recordings_unavailable":0,"recordings_measured":6,"channel_count_min":110,"channel_count_max":110,"sampling_frequency":null,"power_line_frequency":null,"eeg_reference":null,"placement_scheme":null,"sweep_stamps":"{\"enrichment_updated_at\":\"2026-06-17 20:54:58\",\"metadata_updated_at\":\"2026-06-17 20:55:19\",\"archive_checked_at\":\"2026-06-17 20:57:22\",\"zarr_checked_at\":null,\"records_checked_at\":\"2026-06-17 20:56:27\",\"citations_updated_at\":\"2026-09-10 03:00:17\",\"channel_montage_checked_at\":\"2026-06-28 23:00:37\",\"hed_checked_at\":\"2026-06-30 04:30:44\",\"data_checked_at\":null,\"availability_report_at\":\"2026-07-23 01:09:32\",\"signal_defaults_at\":\"2026-09-02 11:48:58\",\"recording_stats_at\":\"2026-09-06 03:00:41\"}","participants":1,"num_citations":15,"latest_version":"v1.0.0","zarr_verify_status":null,"zarr_verified_at":null,"owner_username":"bruaristimunha","owner_github":"bruAristimunha","file_size_formatted":"1.89 GB","zarr_data_failures":null,"zarr_index_url":"https://zarr.nemar.org/nm000251/zarr/index.json","attestation_deposit_type":null,"attestation_key_status":null,"attestation_deidentified":null,"attestation_no_duplicate":null,"attestation_upstream_source":null,"attestation_accepted_at":null}}