{"dataset":{"id":"61257","dataset_id":"on007808","name":"EEG-Speech Brain Decoding Dataset","description":"This dataset comprises EEG and synchronized audio recordings collected for brain-based speech decoding research. It includes overt speech production, auditory listening, and covert speech imagery tasks recorded with multiple EEG acquisition systems (g.tec g.Pangolin, g.tec g.SCARABEO, ANT Neuro eego sports). The dataset supports research on neural correlates of speech production and perception, and brain-computer interface decoding of speech from EEG signals.","owner_user_id":15,"status":"active","github_repo":"nemarDatasets/on007808","concept_doi":"10.82901/nemar.on007808","latest_version_doi":"10.82901/nemar.on007808.v1.0.0","created_at":"2026-06-30 20:31:37","updated_at":"2026-08-18 23:30:31","zenodo_concept_id":null,"is_sandbox":0,"visibility":"public","ezid_status":"public","enrichment_json":"{\n  \"version\": \"2.0\",\n  \"pipeline_stage\": \"validated\",\n  \"title\": \"EEG-Speech Brain Decoding Dataset\",\n  \"description\": \"This dataset comprises EEG and synchronized audio recordings collected for brain-based speech decoding research. It includes overt speech production, auditory listening, and covert speech imagery tasks recorded with multiple EEG acquisition systems (g.tec g.Pangolin, g.tec g.SCARABEO, ANT Neuro eego sports). The dataset supports research on neural correlates of speech production and perception, and brain-computer interface decoding of speech from EEG signals.\",\n  \"methods_description\": \"EEG data were recorded using three different acquisition systems (g.tec g.Pangolin, g.tec g.SCARABEO, ANT Neuro eego sports) across multiple sessions labeled by recording date, with tasks including overt speech production, auditory listening, and covert speech imagery. Audio recordings of vocal production and listening stimuli were collected alongside EEG. Raw EEG data are stored in EDF format, with channel typing detailed in accompanying channels.tsv files and electrode coordinates provided for g.Pangolin sessions.\",\n  \"license\": \"CC0\",\n  \"dataset_type\": \"raw\",\n  \"authors\": {\n    \"Motoshige Sato\": {},\n    \"Ilya Horiguchi\": {},\n    \"Masakazu Inoue\": {\n      \"orcid\": \"0009-0009-2888-936X\"\n    },\n    \"Kenichi Tomeoka\": {},\n    \"Eri Hatakeyama\": {\n      \"orcid\": \"0009-0001-8285-9091\"\n    },\n    \"Yuya Kita\": {},\n    \"Atsushi Yamamoto\": {},\n    \"Ippei Fujisawa\": {},\n    \"Shuntaro Sasai\": {\n      \"orcid\": \"0000-0002-9941-6510\"\n    }\n  },\n  \"keywords\": [\n    {\n      \"term\": \"EEG\"\n    },\n    {\n      \"term\": \"speech production\"\n    },\n    {\n      \"term\": \"Auditory Perception\",\n      \"subject_scheme\": \"MeSH\",\n      \"value_uri\": \"http://id.nlm.nih.gov/mesh/D001307\"\n    },\n    {\n      \"term\": \"brain-computer interface\"\n    },\n    {\n      \"term\": \"covert speech\"\n    },\n    {\n      \"term\": \"speech decoding\"\n    }\n  ],\n  \"related_identifiers\": [\n    {\n      \"identifier\": \"https://github.com/nemarDatasets/on007808\",\n      \"identifier_type\": \"URL\",\n      \"relation_type\": \"IsDescribedBy\"\n    },\n    {\n      \"identifier\": \"https://nemar.org/dataset/on007808\",\n      \"identifier_type\": \"URL\",\n      \"relation_type\": \"IsDescribedBy\"\n    },\n    {\n      \"identifier\": \"10.1088/1741-2552/ae54d0\",\n      \"identifier_type\": \"DOI\",\n      \"relation_type\": \"IsDescribedBy\"\n    },\n    {\n      \"identifier\": \"10.18112/openneuro.ds007808.v1.0.0\",\n      \"identifier_type\": \"DOI\",\n      \"relation_type\": \"IsDerivedFrom\"\n    }\n  ],\n  \"funding_references\": [\n    {\n      \"funder_name\": \"JST\",\n      \"award_number\": \"JPMJMS2012\",\n      \"award_title\": \"Moonshot R&D\"\n    }\n  ],\n  \"resource_type_general\": \"Dataset\",\n  \"resource_type_specific\": \"Neuroimaging Dataset\",\n  \"modalities\": [\n    \"beh\",\n    \"eeg\"\n  ],\n  \"sizes\": [\n    \"1.7 TB (3965 files)\"\n  ],\n  \"formats\": [\n    \".edf\",\n    \".json\",\n    \".md\",\n    \".py\",\n    \".pyc\",\n    \".tsv\",\n    \".wav\",\n    \".yml\"\n  ],\n  \"source_hash\": \"118850a628d010f48a356d29997fdcd5d8e40144bd7f5c8e6c0aeb7c788d33f0\"\n}","last_activity_at":"2026-06-30 20:31:37","source":"openneuro","source_id":"ds007808","subject_count":3,"modalities":"beh,eeg","age_min":22,"age_max":44,"file_size":1691099773494,"total_files":12869,"tasks":"listening,listeningcovert,speechopen","metadata_columns_error":null,"staleness_warn_stage":null,"staleness_admin_notified_at":null,"authors":"Motoshige Sato, Ilya Horiguchi, Masakazu Inoue, Kenichi Tomeoka, Eri Hatakeyama, Yuya Kita, Atsushi Yamamoto, Ippei Fujisawa, Shuntaro Sasai","license":"CC0","readme":"[![DOI](https://img.shields.io/badge/DOI-10.82901%2Fnemar.on007808-blue)](https://doi.org/10.82901/nemar.on007808)\n\n# EEG-Speech Brain Decoding Dataset\n\n## Overview\nThis dataset contains EEG recordings and audio data.\n\n## Sessions\nSessions are labeled by recording date in YYYYMMDD format.\n- Example: `ses-20240401` = recorded on April 1, 2024\n\nMultiple recordings on the same day are distinguished by run numbers:\n- `run-N`: Nth recording of the day\n\n## Tasks\n- **speechopen**: Overt speech production task\n  - Participants vocalize visually presented text\n- **listening**: Auditory listening task\n  - Participants listen to prerecorded speech stimuli\n- **listeningcovert**: Auditory listening followed by covert speech imagery\n\n## EEG Acquisition Devices\nRecordings are split by acquisition label:\n- **acq-pangolin**: g.tec g.Pangolin. Stored channels include 128 EEG channels plus audio monitor, EOG, EMG, trigger, and in 140-channel recordings one auxiliary mastoid/reference channel that is not used for analysis.\n- **acq-scarabeo**: g.tec g.SCARABEO. Stored channels include 64 EEG-related channels, with channels 62 and 63 corresponding to mastoid electrodes, plus audio monitor, EOG, EMG, and trigger channels.\n- **acq-eego**: ANT Neuro eego sports. Stored channels include EEG channels 0..30 and 32..63, channel 31 as miscellaneous, audio monitor, EOG, upper- and lower-lip EMG, miscellaneous auxiliary channels, and trigger.\n\nRun-level `*_channels.tsv` files provide the authoritative channel typing for each recording.\nFor g.Pangolin sessions, `*_electrodes.tsv` and `*_coordsystem.json` files provide 3D coordinates for EEG001..EEG128 from g.tec electrode digitization (`electrodes_uhd.xml`). Electrode coordinate files are not provided for g.SCARABEO or eego sports sessions because verified channel-to-position coordinate mappings are not available in this release.\n\n## File Format Notes\n\n### EEG Data\nRaw EEG data is stored:\n- **Path**: `sub-*/ses-*/eeg/*_eeg.edf`\n- **Note**: EDF files are included as the raw EEG recordings for this BIDS-EEG dataset.\n\n### Behavioral Data (Audio)\nTask audio files are stored in `beh/` directories:\n- **Speech production**: `sub-*/ses-*/beh/*_recording-vocal_beh.wav`\n- **Listening**: `sub-*/ses-*/beh/*_recording-audio_beh.wav`\n- **Note**: Not officially part of BIDS-EEG spec, but included for analysis convenience\n- Excluded in `.bidsignore`\n\n## Directory Structure\n```\ndataset_root/\n├── README                          (this file)\n├── CHANGES                         (version history)\n├── dataset_description.json        (dataset metadata)\n├── participants.tsv                (participant information)\n├── participants.json               (participant column descriptions)\n├── task-speechopen_acq-pangolin_eeg.json               (speech production EEG metadata)\n├── task-speechopen_acq-scarabeo_eeg.json               (speech production EEG metadata)\n├── task-speechopen_acq-eego_eeg.json                   (speech production EEG metadata)\n├── task-listening_acq-pangolin_eeg.json                (listening EEG metadata)\n├── task-speechopen_acq-pangolin_events.json            (speech production events column descriptions)\n├── task-speechopen_acq-pangolin_recording-vocal_beh.json\n├── task-listening_acq-pangolin_recording-audio_beh.json\n├── .bidsignore                     (files to ignore in validation)\n│\n├── code/                           (analysis and preprocessing code)\n│   ├── preprocessing/              (EEG and audio preprocessing)\n│   ├── training/                   (model training scripts)\n│   ├── evaluation/                 (evaluation metrics)\n│   └── bids/                       (BIDS conversion scripts)\n│\n├── sub-01/                         (participant data)\n│   └── ses-YYYYMMDD/              (session by date)\n│       ├── eeg/                    (EEG recordings)\n│       └── beh/                    (behavioral/audio data)\n│\n└── derivatives/                    (processed data)\n    └── pipeline-standard/          (standard preprocessing)\n```\n","bids_version":"1.9.0","sessions_count":312,"publish_date":null,"embedding_dirty":0,"license_tier":"public","zarr_status":"ready","zarr_converted_at":"2026-08-22 23:48:50","zarr_store_count":1973,"zarr_index_etag":"e782fd9b68846e8fa9c8f8f614dde54a","zarr_source_commit":"732fd9bcb7cc5eb4f69c6005dfd4b27a90d4cd01","archive_status":null,"archive_size":null,"archive_retry_count":0,"records_status":null,"archive_skip_reason":"dataset 1575.0 GB exceeds 100.0 GB archive limit; use direct download","zarr_errors":1,"zarr_failure_count":0,"zarr_deterministic":0,"zarr_failed_at":"2026-08-22 23:48:50","num_dataset_citations":1,"num_datapaper_citations":2,"n_channels":128,"electrode_system":"other","has_hed":0,"hed_version":null,"is_exemplar":0,"bytes_present":1691038142989,"data_complete":1,"withdrawn_at":null,"withdrawn_reason":null,"archive_complete":null,"archive_absent_files":null,"archive_declared_files":null,"zarr_pool_breaks":null,"total_recording_duration":3668341,"recording_duration_min":63,"recording_duration_max":10974,"recording_count":1973,"recordings_unavailable":0,"recordings_measured":1973,"channel_count_min":75,"channel_count_max":140,"sampling_frequency":1200,"power_line_frequency":50,"eeg_reference":"FCz","placement_scheme":"Custom left-hemisphere-dominant subset from a high-density g.Pangolin grid","sweep_stamps":"{\"enrichment_updated_at\":\"2026-08-18 23:30:11\",\"metadata_updated_at\":\"2026-08-18 23:30:27\",\"archive_checked_at\":\"2026-06-30 21:14:30\",\"zarr_checked_at\":null,\"records_checked_at\":null,\"citations_updated_at\":\"2026-09-08 03:00:51\",\"channel_montage_checked_at\":null,\"hed_checked_at\":null,\"data_checked_at\":\"2026-09-05 03:00:25\",\"availability_report_at\":\"2026-07-23 01:33:53\",\"recording_stats_at\":\"2026-09-02 11:34:07\",\"signal_defaults_at\":\"2026-09-02 12:58:29\"}","participants":3,"num_citations":3,"latest_version":"v1.0.0","zarr_verify_status":null,"zarr_verified_at":null,"owner_username":"nemarAdmin","owner_github":"nemarAdmin","file_size_formatted":"1.54 TB","zarr_data_failures":null,"zarr_index_url":"https://zarr.nemar.org/on007808/zarr/index.json","attestation_deposit_type":null,"attestation_key_status":null,"attestation_deidentified":null,"attestation_no_duplicate":null,"attestation_upstream_source":null,"attestation_accepted_at":null}}