{"dataset":{"id":"47565","dataset_id":"nm000224","name":"Imagined speech EEG dataset — short and long words (Nguyen et al. 2017)","description":"This dataset comprises preprocessed EEG recordings from 6 healthy participants performing imagined speech tasks, specifically discriminating between short and long words ('cooperate' vs 'in'). Data were acquired at 256 Hz using 64 EEG channels and processed with bandpass filtering (8-70 Hz), notch filtering (60 Hz), and artifact removal. The dataset contains 1,200 trials. Classification results (mean accuracy 73.3±8.9%) reported in the original Nguyen et al. 2017 study are provided for reference.","owner_user_id":19,"status":"active","github_repo":"nemarDatasets/nm000224","concept_doi":"10.82901/nemar.nm000224","latest_version_doi":"10.82901/nemar.nm000224.v1.0.3","created_at":"2026-06-19 23:06:10","updated_at":"2026-08-18 18:18:14","zenodo_concept_id":null,"is_sandbox":0,"visibility":"public","ezid_status":"public","enrichment_json":"{\n  \"version\": \"2.0\",\n  \"pipeline_stage\": \"enriched\",\n  \"title\": \"Imagined speech EEG dataset — short and long words (Nguyen et al. 2017)\",\n  \"description\": \"This dataset comprises preprocessed EEG recordings from 6 healthy participants performing imagined speech tasks, specifically discriminating between short and long words ('cooperate' vs 'in'). Data were acquired at 256 Hz using 64 EEG channels and processed with bandpass filtering (8-70 Hz), notch filtering (60 Hz), and artifact removal. The dataset contains 1,200 trials. Classification results (mean accuracy 73.3±8.9%) reported in the original Nguyen et al. 2017 study are provided for reference.\",\n  \"methods_description\": \"EEG data were acquired using a BrainProducts ActiCHamp system at 256 Hz sampling rate with 64 channels (60 EEG, 4 EOG) in a standard 10-20 montage. Participants performed imagined speech tasks cued by auditory beeps (5 beeps at 1.4 s rhythm) and visual cues, with 5-second analysis windows split into 3 overlapping 2-second epochs. Preprocessing included 5th-order Butterworth bandpass filtering (8-70 Hz), 60 Hz notch filtering, and adaptive EOG artifact removal. Analysis employed Riemannian manifold and relevance vector machine approaches.\",\n  \"license\": \"other-open\",\n  \"dataset_type\": \"derivative\",\n  \"authors\": {\n    \"Chuong H. Nguyen\": {},\n    \"George K. Karavas\": {},\n    \"Panagiotis Artemiadis\": {}\n  },\n  \"keywords\": [\n    {\n      \"term\": \"imagined speech\"\n    },\n    {\n      \"term\": \"EEG\"\n    },\n    {\n      \"term\": \"Brain-Computer Interfaces\",\n      \"subject_scheme\": \"MeSH\",\n      \"scheme_uri\": \"https://id.nlm.nih.gov/mesh/\",\n      \"value_uri\": \"http://id.nlm.nih.gov/mesh/D062207\"\n    },\n    {\n      \"term\": \"Riemannian manifold\"\n    },\n    {\n      \"term\": \"covariance matrix\"\n    },\n    {\n      \"term\": \"speech imagery\"\n    },\n    {\n      \"term\": \"relevance vector machines\"\n    },\n    {\n      \"term\": \"short vs long words\"\n    }\n  ],\n  \"related_identifiers\": [\n    {\n      \"identifier\": \"10.1088/1741-2552/aa8235\",\n      \"identifier_type\": \"DOI\",\n      \"relation_type\": \"IsDerivedFrom\"\n    },\n    {\n      \"identifier\": \"https://github.com/nemarDatasets/nm000224\",\n      \"identifier_type\": \"URL\",\n      \"relation_type\": \"IsDescribedBy\"\n    },\n    {\n      \"identifier\": \"10.21105/joss.01896\",\n      \"identifier_type\": \"DOI\",\n      \"relation_type\": \"References\"\n    },\n    {\n      \"identifier\": \"10.1038/s41597-019-0104-8\",\n      \"identifier_type\": \"DOI\",\n      \"relation_type\": \"References\"\n    },\n    {\n      \"identifier\": \"https://nemar.org/dataset/nm000224\",\n      \"identifier_type\": \"URL\",\n      \"relation_type\": \"IsDescribedBy\"\n    }\n  ],\n  \"resource_type_general\": \"Dataset\",\n  \"resource_type_specific\": \"EEG Dataset\",\n  \"modalities\": [\n    \"eeg\"\n  ],\n  \"sizes\": [\n    \"1.2 GB (13 files)\"\n  ],\n  \"formats\": [\n    \".edf\",\n    \".json\",\n    \".mat\",\n    \".md\",\n    \".tsv\",\n    \".yaml\",\n    \".yml\"\n  ],\n  \"source_hash\": \"8188404fc11bee1bcb394e8e232c12605e572ddf4d123132e687fb4b035b7c2f\"\n}","last_activity_at":"2026-08-16 13:33:38","source":null,"source_id":null,"subject_count":6,"modalities":"eeg","age_min":null,"age_max":null,"file_size":1229895222,"total_files":83,"tasks":"imagery","metadata_columns_error":null,"staleness_warn_stage":null,"staleness_admin_notified_at":null,"authors":"Chuong H. Nguyen, George K. Karavas, Panagiotis Artemiadis","license":"other-open","readme":"[![DOI](https://img.shields.io/badge/DOI-10.82901%2Fnemar.nm000224-blue)](https://doi.org/10.82901/nemar.nm000224)\n\n# Imagined speech EEG dataset — short and long words (Nguyen et al. 2017)\n\n## Overview\n\nThis dataset comprises preprocessed EEG recordings from 6 healthy participants performing imagined speech tasks, specifically discriminating between short and long words ('cooperate' vs 'in'). Data were acquired at 256 Hz using 64 EEG channels and processed with bandpass filtering (8-70 Hz), notch filtering (60 Hz), and artifact removal. The dataset contains 3,600 trials analyzed using Riemannian manifold and relevance vector machine approaches for brain-computer interface applications, achieving mean classification accuracy of 73.3±8.9%.\n\n## Dataset Summary\n\n| Property | Value |\n|---|---|\n| Subjects | 6 |\n| Channels | 64 |\n| Classes | 2 |\n| Trial length | 5 s |\n| Sampling frequency | 256 Hz |\n| Sessions | 1 |\n| Total trials | 1200 |\n| Paradigm | MotorImagery |\n\n## Data Collection Methods\n\nEEG data were acquired using a BrainProducts ActiCHamp system at 256 Hz sampling rate with 64 channels (60 EEG, 4 EOG) in a standard 1020 montage. Participants performed imagined speech tasks cued by auditory beeps (5 beeps at 1.4 s rhythm) and visual cues, with 5-second analysis windows split into 3 overlapping 2-second epochs. Preprocessing included 5th-order Butterworth bandpass filtering (8-70 Hz), 60 Hz notch filtering, and adaptive EOG artifact removal.\n\n## How to Access via MOABB\n\nInstall MOABB and load this dataset directly:\n\n```python\nfrom moabb.datasets import Nguyen2017_SL\nfrom moabb.paradigms import MotorImagery\nparadigm = MotorImagery()\n\ndataset = Nguyen2017_SL()\nX, y, metadata = paradigm.get_data(dataset)\n```\n\nFor more details see the [MOABB documentation](https://moabb.neurotechx.com/) and the\n[MOABB dataset page](https://moabb.neurotechx.com/docs/generated/moabb.datasets.Nguyen2017_SL.html).\n\n## Citation\n\nIf you use this dataset please cite the primary publication:\n\n> DOI: [10.1088/1741-2552/aa8235](https://doi.org/10.1088/1741-2552/aa8235)\n\n## NEMAR / MOABB Benchmark Collection\n\nThis BIDS-formatted dataset was converted from the original data using the\n[MOABB](https://moabb.neurotechx.com/) pipeline and re-hosted on\n[NEMAR](https://nemar.org/) as part of the MOABB benchmark collection.\nThe original data and license terms apply — see `dataset_description.json` for details.\n","bids_version":"1.9.0","sessions_count":1,"publish_date":null,"embedding_dirty":0,"license_tier":"unknown","zarr_status":"ready","zarr_converted_at":"2026-08-22 07:43:54","zarr_store_count":6,"zarr_index_etag":"4af56e61365490f6ad48f9d62c648beb","zarr_source_commit":"fd9d556497f1204f5a5d5bac5c9828f7a01f0aae","archive_status":"ready","archive_size":1165416201,"archive_retry_count":0,"records_status":"ready","archive_skip_reason":null,"zarr_errors":0,"zarr_failure_count":0,"zarr_deterministic":0,"zarr_failed_at":null,"num_dataset_citations":0,"num_datapaper_citations":98,"n_channels":60,"electrode_system":"10-10","has_hed":1,"hed_version":"8.4.0","is_exemplar":0,"bytes_present":1229713657,"data_complete":1,"withdrawn_at":null,"withdrawn_reason":null,"archive_complete":1,"archive_absent_files":0,"archive_declared_files":83,"zarr_pool_breaks":null,"total_recording_duration":8030,"recording_duration_min":1147,"recording_duration_max":1434,"recording_count":6,"recordings_unavailable":0,"recordings_measured":6,"channel_count_min":60,"channel_count_max":60,"sampling_frequency":256,"power_line_frequency":60,"eeg_reference":null,"placement_scheme":"10-20 system","sweep_stamps":"{\"enrichment_updated_at\":\"2026-08-18 18:15:12\",\"metadata_updated_at\":\"2026-08-18 18:18:12\",\"archive_checked_at\":\"2026-08-18 18:21:07\",\"zarr_checked_at\":null,\"records_checked_at\":\"2026-08-18 18:23:19\",\"citations_updated_at\":\"2026-09-08 03:00:46\",\"channel_montage_checked_at\":\"2026-06-28 22:58:42\",\"hed_checked_at\":\"2026-06-30 04:28:46\",\"data_checked_at\":null,\"availability_report_at\":\"2026-08-21 03:01:16\",\"recording_stats_at\":\"2026-09-02 11:31:56\",\"signal_defaults_at\":\"2026-09-02 11:45:55\"}","participants":6,"num_citations":98,"latest_version":"v1.0.3","zarr_verify_status":null,"zarr_verified_at":null,"owner_username":"bruaristimunha","owner_github":"bruAristimunha","file_size_formatted":"1.15 GB","zarr_data_failures":null,"zarr_index_url":"https://zarr.nemar.org/nm000224/zarr/index.json","attestation_deposit_type":null,"attestation_key_status":null,"attestation_deidentified":null,"attestation_no_duplicate":null,"attestation_upstream_source":null,"attestation_accepted_at":null}}