{"dataset":{"id":"47102","dataset_id":"nm000174","name":"Imagined Speech EEG dataset comparing paradigm designs (Aguilera-Rodriguez et al. 2025)","description":"An EEG dataset of imagined speech from 15 healthy participants comparing traditional cue-based and gamified paradigm designs for brain-computer interface applications. The dataset comprises 1,800 trials of four Spanish directional commands recorded at 500 Hz from 24 electrodes using the mBrainTrain Smarting system. This derivative dataset enables systematic evaluation of paradigm design effects on imagined speech decoding performance.","owner_user_id":19,"status":"active","github_repo":"nemarDatasets/nm000174","concept_doi":"10.82901/nemar.nm000174","latest_version_doi":"10.82901/nemar.nm000174.v1.0.3","created_at":"2026-06-19 17:07:41","updated_at":"2026-08-18 18:15:14","zenodo_concept_id":null,"is_sandbox":0,"visibility":"public","ezid_status":"public","enrichment_json":"{\n  \"version\": \"2.0\",\n  \"pipeline_stage\": \"enriched\",\n  \"title\": \"Imagined Speech EEG dataset comparing paradigm designs (Aguilera-Rodriguez et al. 2025)\",\n  \"description\": \"An EEG dataset of imagined speech from 15 healthy participants comparing traditional cue-based and gamified paradigm designs for brain-computer interface applications. The dataset comprises 1,800 trials of four Spanish directional commands recorded at 500 Hz from 24 electrodes using the mBrainTrain Smarting system. This derivative dataset enables systematic evaluation of paradigm design effects on imagined speech decoding performance.\",\n  \"methods_description\": \"EEG data were acquired from 15 healthy participants (8 male, 7 female, ages 18-27) using a 24-channel mBrainTrain Smarting system at 500 Hz sampling rate with FCz reference and Fpz ground. Two experimental paradigms were compared: (1) traditional cue-based with visual and auditory cues (5 beeps at 1.4s rhythm) followed by 7 imagined speech repetitions, with the last 3 repetitions extracted for analysis, and (2) gamified paradigm using a Pac-Man maze interface with 1.4s periods for movement decision, imagined speech, and vocalized speech. Each participant completed 1 session with 120 trials per session (30 per class), yielding 450 trials per class. The final dataset contains the last 3 imagined speech repetitions per trial from the traditional paradigm.\",\n  \"license\": \"CC-BY-NC-ND-4.0\",\n  \"dataset_type\": \"derivative\",\n  \"authors\": {\n    \"Edgar Aguilera-Rodriguez\": {},\n    \"Alma Cuevas-Romero\": {},\n    \"Santiago Mendoza-Franco\": {\n      \"orcid\": \"0009-0005-7126-6636\"\n    },\n    \"Jonathan Wornovitzky-Green\": {},\n    \"Eduardo Rivera-Cerros\": {},\n    \"David Villanueva-Cazares\": {},\n    \"Luis Alberto Munoz-Ubando\": {},\n    \"David Ibarra-Zarate\": {},\n    \"Luz Maria Alonso-Valerdi\": {\n      \"orcid\": \"0000-0002-2256-2958\"\n    }\n  },\n  \"keywords\": [\n    {\n      \"term\": \"imagined speech\"\n    },\n    {\n      \"term\": \"EEG\"\n    },\n    {\n      \"term\": \"Brain-Computer Interfaces\",\n      \"subject_scheme\": \"MeSH\",\n      \"value_uri\": \"http://id.nlm.nih.gov/mesh/D062207\"\n    },\n    {\n      \"term\": \"motor imagery\",\n      \"subject_scheme\": \"MeSH\",\n      \"value_uri\": \"http://id.nlm.nih.gov/mesh/D019157\"\n    },\n    {\n      \"term\": \"gamified paradigm\"\n    }\n  ],\n  \"related_identifiers\": [\n    {\n      \"identifier\": \"10.1038/s41597-025-05926-5\",\n      \"identifier_type\": \"DOI\",\n      \"relation_type\": \"IsDerivedFrom\"\n    },\n    {\n      \"identifier\": \"https://github.com/nemarDatasets/nm000174\",\n      \"identifier_type\": \"URL\",\n      \"relation_type\": \"IsDescribedBy\"\n    },\n    {\n      \"identifier\": \"10.21105/joss.01896\",\n      \"identifier_type\": \"DOI\",\n      \"relation_type\": \"References\"\n    },\n    {\n      \"identifier\": \"10.1038/s41597-019-0104-8\",\n      \"identifier_type\": \"DOI\",\n      \"relation_type\": \"References\"\n    },\n    {\n      \"identifier\": \"https://nemar.org/dataset/nm000174\",\n      \"identifier_type\": \"URL\",\n      \"relation_type\": \"IsDescribedBy\"\n    }\n  ],\n  \"resource_type_general\": \"Dataset\",\n  \"resource_type_specific\": \"EEG Dataset\",\n  \"modalities\": [\n    \"eeg\"\n  ],\n  \"sizes\": [\n    \"1.6 GB (91 files)\"\n  ],\n  \"formats\": [\n    \".edf\",\n    \".eeg\",\n    \".json\",\n    \".md\",\n    \".tsv\",\n    \".vhdr\",\n    \".vmrk\",\n    \".xdf\",\n    \".yaml\",\n    \".yml\"\n  ],\n  \"source_hash\": \"1a394a19b137ba3158f9f6220ffc32b33d8cf50e3274d5afa1326e1286d52a4b\"\n}","last_activity_at":"2026-08-16 13:27:34","source":null,"source_id":null,"subject_count":15,"modalities":"eeg","age_min":null,"age_max":null,"file_size":1633047048,"total_files":431,"tasks":"imagery","metadata_columns_error":null,"staleness_warn_stage":null,"staleness_admin_notified_at":null,"authors":"Edgar Aguilera-Rodriguez, Alma Cuevas-Romero, Santiago Mendoza-Franco, Jonathan Wornovitzky-Green, Eduardo Rivera-Cerros, David Villanueva-Cazares, Luis Alberto Munoz-Ubando, David Ibarra-Zarate, Luz Maria Alonso-Valerdi","license":"CC-BY-NC-ND-4.0","readme":"[![DOI](https://img.shields.io/badge/DOI-10.82901%2Fnemar.nm000174-blue)](https://doi.org/10.82901/nemar.nm000174)\n\n# Imagined Speech EEG dataset comparing paradigm designs (Aguilera-Rodriguez et al. 2025)\n\n## Overview\n\nAn EEG dataset of imagined speech from 15 healthy participants comparing traditional cue-based and gamified (Pac-Man maze) paradigms for brain-computer interface applications. The dataset comprises 1,800 trials across four Spanish directional commands (avanzar, retroceder, derecha, izquierda) recorded at 500 Hz from 24 electrodes using the mBrainTrain Smarting system. This derivative dataset enables systematic evaluation of paradigm design effects on imagined speech decoding performance.\n\n## Dataset Summary\n\n| Property | Value |\n|---|---|\n| Subjects | 15 |\n| Channels | 24 |\n| Classes | 4 |\n| Trial length | 4 s |\n| Sampling frequency | 500 Hz |\n| Sessions | 1 |\n| Total trials | 1800 |\n| Paradigm | MotorImagery |\n\n## Data Collection Methods\n\nEEG data were acquired from 15 healthy participants (8 male, 7 female, ages 18-27) using a 24-channel mBrainTrain Smarting system at 500 Hz sampling rate with FCz reference and Fpz ground. Two experimental paradigms were compared: (1) traditional cue-based with visual and auditory cues (5 beeps at 1.4s rhythm) followed by 7 imagined speech repetitions, with the last 3 repetitions extracted for analysis, and (2) gamified paradigm using a Pac-Man maze interface with 1.4s periods for movement decision, imagined speech, and vocalized speech. Each participant completed 2 sessions with 120 trials per session (30 per class), yielding 450 trials per class across the dataset. The final dataset contains the last 3 imagined speech repetitions per trial from the traditional paradigm.\n\n## How to Access via MOABB\n\nInstall MOABB and load this dataset directly:\n\n```python\nfrom moabb.datasets import AguileraRodriguez2025\nfrom moabb.paradigms import MotorImagery\nparadigm = MotorImagery()\n\ndataset = AguileraRodriguez2025()\nX, y, metadata = paradigm.get_data(dataset)\n```\n\nFor more details see the [MOABB documentation](https://moabb.neurotechx.com/) and the\n[MOABB dataset page](https://moabb.neurotechx.com/docs/generated/moabb.datasets.AguileraRodriguez2025.html).\n\n## Citation\n\nIf you use this dataset please cite the primary publication:\n\n> DOI: [10.1038/s41597-025-05926-5](https://doi.org/10.1038/s41597-025-05926-5)\n\n## NEMAR / MOABB Benchmark Collection\n\nThis BIDS-formatted dataset was converted from the original data using the\n[MOABB](https://moabb.neurotechx.com/) pipeline and re-hosted on\n[NEMAR](https://nemar.org/) as part of the MOABB benchmark collection.\nThe original data and license terms apply — see `dataset_description.json` for details.\n","bids_version":"1.9.0","sessions_count":2,"publish_date":null,"embedding_dirty":0,"license_tier":"noderiv","zarr_status":"ready","zarr_converted_at":"2026-08-22 06:58:00","zarr_store_count":30,"zarr_index_etag":"56174a888100d1ad8bb3f150acc85583","zarr_source_commit":"52a25534108a1bff834e187cf6693f3495920356","archive_status":"ready","archive_size":1380210901,"archive_retry_count":0,"records_status":"ready","archive_skip_reason":null,"zarr_errors":0,"zarr_failure_count":0,"zarr_deterministic":0,"zarr_failed_at":null,"num_dataset_citations":0,"num_datapaper_citations":1,"n_channels":24,"electrode_system":"10-10","has_hed":1,"hed_version":"8.4.0","is_exemplar":0,"bytes_present":1632582583,"data_complete":1,"withdrawn_at":null,"withdrawn_reason":null,"archive_complete":1,"archive_absent_files":0,"archive_declared_files":431,"zarr_pool_breaks":null,"total_recording_duration":18106.596,"recording_duration_min":492,"recording_duration_max":720.18,"recording_count":30,"recordings_unavailable":0,"recordings_measured":30,"channel_count_min":24,"channel_count_max":24,"sampling_frequency":500,"power_line_frequency":60,"eeg_reference":"FCz","placement_scheme":"10-05 system","sweep_stamps":"{\"enrichment_updated_at\":\"2026-08-18 18:12:59\",\"metadata_updated_at\":\"2026-08-18 18:15:11\",\"archive_checked_at\":\"2026-08-18 18:22:15\",\"zarr_checked_at\":null,\"records_checked_at\":\"2026-08-18 18:21:37\",\"citations_updated_at\":\"2026-09-08 03:00:52\",\"channel_montage_checked_at\":\"2026-06-28 22:55:19\",\"hed_checked_at\":\"2026-06-30 04:15:12\",\"data_checked_at\":null,\"availability_report_at\":\"2026-08-20 03:00:34\",\"recording_stats_at\":\"2026-09-02 11:31:46\",\"signal_defaults_at\":\"2026-09-02 11:41:29\"}","participants":15,"num_citations":1,"latest_version":"v1.0.3","zarr_verify_status":null,"zarr_verified_at":null,"owner_username":"bruaristimunha","owner_github":"bruAristimunha","file_size_formatted":"1.52 GB","zarr_data_failures":null,"zarr_index_url":"https://zarr.nemar.org/nm000174/zarr/index.json","attestation_deposit_type":null,"attestation_key_status":null,"attestation_deidentified":null,"attestation_no_duplicate":null,"attestation_upstream_source":null,"attestation_accepted_at":null}}