{"dataset":{"id":"47570","dataset_id":"nm000261","name":"Imagined speech EEG dataset — vowels condition (Nguyen et al. 2017)","description":"This dataset comprises preprocessed EEG recordings from 8 healthy subjects performing imagined speech tasks involving three vowel phonemes (a, i, u). Participants received auditory and visual cues to imagine speaking each vowel, with 64-channel EEG data acquired at 256 Hz. The dataset includes 2,400 trials generating 7,200 overlapping 2-second epochs (8 subjects × 300 trials × 3 overlapping epochs per trial) analyzed using Riemannian manifold-based feature extraction and relevance vector machine classification for brain-computer interface applications, achieving approximately 49% mean accuracy across a 10-fold cross-validation scheme.","owner_user_id":19,"status":"active","github_repo":"nemarDatasets/nm000261","concept_doi":"10.82901/nemar.nm000261","latest_version_doi":"10.82901/nemar.nm000261.v1.0.3","created_at":"2026-06-19 23:14:23","updated_at":"2026-08-18 18:19:37","zenodo_concept_id":null,"is_sandbox":0,"visibility":"public","ezid_status":"public","enrichment_json":"{\n  \"version\": \"2.0\",\n  \"pipeline_stage\": \"enriched\",\n  \"title\": \"Imagined speech EEG dataset — vowels condition (Nguyen et al. 2017)\",\n  \"description\": \"This dataset comprises preprocessed EEG recordings from 8 healthy subjects performing imagined speech tasks involving three vowel phonemes (a, i, u). Participants received auditory and visual cues to imagine speaking each vowel, with 64-channel EEG data acquired at 256 Hz. The dataset includes 2,400 trials generating 7,200 overlapping 2-second epochs (8 subjects × 300 trials × 3 overlapping epochs per trial) analyzed using Riemannian manifold-based feature extraction and relevance vector machine classification for brain-computer interface applications, achieving approximately 49% mean accuracy across a 10-fold cross-validation scheme.\",\n  \"methods_description\": \"EEG data were acquired using a BrainProducts ActiCHamp system with 64 channels (60 EEG, 4 EOG) at 256 Hz sampling rate using a standard 10-20 montage. The experimental paradigm consisted of auditory beeps at 1.0 s intervals with visual cues, during which subjects imagined speaking designated vowels. Preprocessing included bandpass filtering (8-70 Hz, 5th order Butterworth), 60 Hz notch filtering, and adaptive EOG artifact removal. Feature extraction employed Riemannian tangent space mapping with common spatial pattern spatial filtering across mu (8-13 Hz), beta (13-30 Hz), and gamma (30-70 Hz) frequency bands. Classification performance was evaluated using a 10-fold cross-validation scheme with relevance vector machines, achieving approximately 49% mean accuracy.\",\n  \"license\": \"other-open\",\n  \"dataset_type\": \"derivative\",\n  \"authors\": {\n    \"Chuong H. Nguyen\": {},\n    \"George K. Karavas\": {},\n    \"Panagiotis Artemiadis\": {}\n  },\n  \"keywords\": [\n    {\n      \"term\": \"imagined speech\"\n    },\n    {\n      \"term\": \"EEG\"\n    },\n    {\n      \"term\": \"brain-computer interface\"\n    },\n    {\n      \"term\": \"motor imagery\"\n    },\n    {\n      \"term\": \"Riemannian manifold\"\n    },\n    {\n      \"term\": \"covariance matrix\"\n    },\n    {\n      \"term\": \"vowels\"\n    },\n    {\n      \"term\": \"relevance vector machines\"\n    }\n  ],\n  \"related_identifiers\": [\n    {\n      \"identifier\": \"10.1088/1741-2552/aa8235\",\n      \"identifier_type\": \"DOI\",\n      \"relation_type\": \"IsDerivedFrom\"\n    },\n    {\n      \"identifier\": \"https://github.com/nemarDatasets/nm000261\",\n      \"identifier_type\": \"URL\",\n      \"relation_type\": \"IsDescribedBy\"\n    },\n    {\n      \"identifier\": \"10.21105/joss.01896\",\n      \"identifier_type\": \"DOI\",\n      \"relation_type\": \"References\"\n    },\n    {\n      \"identifier\": \"10.1038/s41597-019-0104-8\",\n      \"identifier_type\": \"DOI\",\n      \"relation_type\": \"References\"\n    },\n    {\n      \"identifier\": \"https://nemar.org/dataset/nm000261\",\n      \"identifier_type\": \"URL\",\n      \"relation_type\": \"IsDescribedBy\"\n    }\n  ],\n  \"resource_type_general\": \"Dataset\",\n  \"resource_type_specific\": \"EEG Dataset\",\n  \"modalities\": [\n    \"eeg\"\n  ],\n  \"sizes\": [\n    \"2.8 GB (17 files)\"\n  ],\n  \"formats\": [\n    \".edf\",\n    \".json\",\n    \".mat\",\n    \".md\",\n    \".tsv\",\n    \".yaml\",\n    \".yml\"\n  ],\n  \"source_hash\": \"909fd8b8b6333cd18cba4b7f323f7366d2f3acb08cfb82304c98e91290288052\"\n}","last_activity_at":"2026-08-16 13:36:39","source":null,"source_id":null,"subject_count":8,"modalities":"eeg","age_min":null,"age_max":null,"file_size":2768975454,"total_files":107,"tasks":"imagery","metadata_columns_error":null,"staleness_warn_stage":null,"staleness_admin_notified_at":null,"authors":"Chuong H. Nguyen, George K. Karavas, Panagiotis Artemiadis","license":"other-open","readme":"[![DOI](https://img.shields.io/badge/DOI-10.82901%2Fnemar.nm000261-blue)](https://doi.org/10.82901/nemar.nm000261)\n\n# Imagined speech EEG dataset — vowels condition (Nguyen et al. 2017)\n\n## Overview\n\nThis dataset comprises preprocessed EEG recordings from 8 healthy subjects performing imagined speech tasks involving three vowel phonemes (a, i, u). Participants received auditory and visual cues to imagine speaking each vowel, with 64-channel EEG data acquired at 256 Hz. The dataset includes 7,200 trials (8 subjects × 300 trials × 3 overlapping 2-second epochs per trial) analyzed using Riemannian manifold-based feature extraction and relevance vector machine classification, achieving approximately 49% mean accuracy across a 10-fold cross-validation scheme.\n\n## Dataset Summary\n\n| Property | Value |\n|---|---|\n| Subjects | 8 |\n| Channels | 64 |\n| Classes | 3 |\n| Trial length | 5 s |\n| Sampling frequency | 256 Hz |\n| Sessions | 1 |\n| Total trials | 2400 |\n| Paradigm | MotorImagery |\n\n## Data Collection Methods\n\nEEG data were acquired using a BrainProducts ActiCHamp system with 64 channels (60 EEG, 4 EOG) at 256 Hz sampling rate using a standard 10-20 montage. The experimental paradigm consisted of auditory beeps at 1.0 s intervals with visual cues, during which subjects imagined speaking designated vowels. Preprocessing included bandpass filtering (8-70 Hz, 5th order Butterworth), 60 Hz notch filtering, and adaptive EOG artifact removal. Feature extraction employed Riemannian tangent space mapping with common spatial pattern spatial filtering across mu (8-13 Hz), beta (13-30 Hz), and gamma (30-70 Hz) frequency bands.\n\n## How to Access via MOABB\n\nInstall MOABB and load this dataset directly:\n\n```python\nfrom moabb.datasets import Nguyen2017_V\nfrom moabb.paradigms import MotorImagery\nparadigm = MotorImagery()\n\ndataset = Nguyen2017_V()\nX, y, metadata = paradigm.get_data(dataset)\n```\n\nFor more details see the [MOABB documentation](https://moabb.neurotechx.com/) and the\n[MOABB dataset page](https://moabb.neurotechx.com/docs/generated/moabb.datasets.Nguyen2017_V.html).\n\n## Citation\n\nIf you use this dataset please cite the primary publication:\n\n> DOI: [10.1088/1741-2552/aa8235](https://doi.org/10.1088/1741-2552/aa8235)\n\n## NEMAR / MOABB Benchmark Collection\n\nThis BIDS-formatted dataset was converted from the original data using the\n[MOABB](https://moabb.neurotechx.com/) pipeline and re-hosted on\n[NEMAR](https://nemar.org/) as part of the MOABB benchmark collection.\nThe original data and license terms apply — see `dataset_description.json` for details.\n","bids_version":"1.9.0","sessions_count":1,"publish_date":null,"embedding_dirty":0,"license_tier":"unknown","zarr_status":"ready","zarr_converted_at":"2026-08-22 07:37:54","zarr_store_count":8,"zarr_index_etag":"a8552ef6ef3f250bcdc665e4ce0601a1","zarr_source_commit":"e32941842cad034213caa96bfac7ab34fc48452a","archive_status":"ready","archive_size":2615813067,"archive_retry_count":0,"records_status":"ready","archive_skip_reason":null,"zarr_errors":0,"zarr_failure_count":0,"zarr_deterministic":0,"zarr_failed_at":null,"num_dataset_citations":0,"num_datapaper_citations":96,"n_channels":60,"electrode_system":"10-10","has_hed":1,"hed_version":"8.4.0","is_exemplar":0,"bytes_present":2768635143,"data_complete":1,"withdrawn_at":null,"withdrawn_reason":null,"archive_complete":1,"archive_absent_files":0,"archive_declared_files":107,"zarr_pool_breaks":null,"total_recording_duration":17208,"recording_duration_min":2151,"recording_duration_max":2151,"recording_count":8,"recordings_unavailable":0,"recordings_measured":8,"channel_count_min":60,"channel_count_max":60,"sampling_frequency":256,"power_line_frequency":60,"eeg_reference":null,"placement_scheme":"10-20 system","sweep_stamps":"{\"enrichment_updated_at\":\"2026-08-18 18:16:33\",\"metadata_updated_at\":\"2026-08-18 18:19:36\",\"archive_checked_at\":\"2026-08-18 18:25:02\",\"zarr_checked_at\":null,\"records_checked_at\":\"2026-08-18 18:23:46\",\"citations_updated_at\":\"2026-09-08 03:00:46\",\"channel_montage_checked_at\":\"2026-06-28 23:01:36\",\"hed_checked_at\":\"2026-06-30 04:31:49\",\"data_checked_at\":null,\"availability_report_at\":\"2026-08-22 03:01:14\",\"recording_stats_at\":\"2026-09-02 11:32:05\",\"signal_defaults_at\":\"2026-09-02 11:49:43\"}","participants":8,"num_citations":96,"latest_version":"v1.0.3","zarr_verify_status":null,"zarr_verified_at":null,"owner_username":"bruaristimunha","owner_github":"bruAristimunha","file_size_formatted":"2.58 GB","zarr_data_failures":null,"zarr_index_url":"https://zarr.nemar.org/nm000261/zarr/index.json","attestation_deposit_type":null,"attestation_key_status":null,"attestation_deidentified":null,"attestation_no_duplicate":null,"attestation_upstream_source":null,"attestation_accepted_at":null}}