{"dataset":{"id":"61247","dataset_id":"on007629","name":"ROAMM","description":"ROAMM is a large-scale multimodal dataset combining simultaneous EEG and eye-tracking recordings collected during naturalistic reading, with span-level mind-wandering annotations from 44 participants. It provides a benchmark for mind-wandering detection and EEG-to-text decoding, supporting research on attention-related degradation in language decoding from brain activity during naturalistic reading.","owner_user_id":15,"status":"active","github_repo":"nemarDatasets/on007629","concept_doi":"10.82901/nemar.on007629","latest_version_doi":"10.82901/nemar.on007629.v1.0.0","created_at":"2026-06-30 16:01:37","updated_at":"2026-08-18 23:39:36","zenodo_concept_id":null,"is_sandbox":0,"visibility":"public","ezid_status":"public","enrichment_json":"{\n  \"version\": \"2.0\",\n  \"pipeline_stage\": \"validated\",\n  \"title\": \"ROAMM\",\n  \"description\": \"ROAMM is a large-scale multimodal dataset combining simultaneous EEG and eye-tracking recordings collected during naturalistic reading, with span-level mind-wandering annotations from 44 participants. It provides a benchmark for mind-wandering detection and EEG-to-text decoding, supporting research on attention-related degradation in language decoding from brain activity during naturalistic reading.\",\n  \"methods_description\": \"Data were collected using a BioSemi ActiveTwo 64-channel EEG system with simultaneous eye-tracking via an SR Research EyeLink 1000 Plus, during a naturalistic reading task (ReMind) involving standardized articles and retrospective self-report of mind-wandering, along with page-level multiple-choice reading comprehension assessments.\",\n  \"license\": \"CC0\",\n  \"dataset_type\": \"raw\",\n  \"authors\": {\n    \"Haorui Sun\": {},\n    \"Ardyn Vivienne Olszko\": {},\n    \"Niharika Singh\": {},\n    \"David C. Jangraw\": {}\n  },\n  \"keywords\": [\n    {\n      \"term\": \"EEG\"\n    },\n    {\n      \"term\": \"Eye-Tracking\"\n    },\n    {\n      \"term\": \"mind-wandering\"\n    },\n    {\n      \"term\": \"naturalistic reading\"\n    },\n    {\n      \"term\": \"EEG-to-text decoding\"\n    },\n    {\n      \"term\": \"attention\"\n    }\n  ],\n  \"related_identifiers\": [\n    {\n      \"identifier\": \"https://github.com/nemarDatasets/on007629\",\n      \"identifier_type\": \"URL\",\n      \"relation_type\": \"IsDescribedBy\"\n    },\n    {\n      \"identifier\": \"https://nemar.org/dataset/on007629\",\n      \"identifier_type\": \"URL\",\n      \"relation_type\": \"IsDescribedBy\"\n    },\n    {\n      \"identifier\": \"10.18112/openneuro.ds007629.v1.1.0\",\n      \"identifier_type\": \"DOI\",\n      \"relation_type\": \"IsDerivedFrom\"\n    }\n  ],\n  \"resource_type_general\": \"Dataset\",\n  \"resource_type_specific\": \"Neuroimaging Dataset\",\n  \"modalities\": [\n    \"beh\",\n    \"eeg\"\n  ],\n  \"sizes\": [\n    \"50.0 GB (246 files)\"\n  ],\n  \"formats\": [\n    \".csv\",\n    \".eeg\",\n    \".gz\",\n    \".json\",\n    \".md\",\n    \".pkl\",\n    \".tsv\",\n    \".vhdr\",\n    \".vmrk\",\n    \".yml\"\n  ],\n  \"source_hash\": \"4c5349975ab0a30a08c89668d54b1d1eaf471ba076be28f704855285280fa156\"\n}","last_activity_at":"2026-06-30 16:01:37","source":"openneuro","source_id":"ds007629","subject_count":1,"modalities":"beh,eeg","age_min":18,"age_max":64,"file_size":50025343803,"total_files":273,"tasks":"ReMind","metadata_columns_error":null,"staleness_warn_stage":null,"staleness_admin_notified_at":null,"authors":"Haorui Sun, Ardyn Vivienne Olszko, Niharika Singh, David C. Jangraw","license":"CC0","readme":"[![DOI](https://img.shields.io/badge/DOI-10.82901%2Fnemar.on007629-blue)](https://doi.org/10.82901/nemar.on007629)\n\n# ROAMM: Reading Observed At Mindless Moments\n\n**ROAMM** is a large-scale multimodal dataset featuring simultaneous **EEG and eye-tracking** data collected during naturalistic reading with **span-level mind-wandering annotations**. ROAMM provides a benchmark dataset for MW detection and EEG-to-text decoding tasks, and enables the study of attention-related degradation in language decoding from brain activity in naturalistic reading.\n\n## Dataset Status\n* **Synchronized ML Dataset:** For researchers looking for the pre-processed, synchronized EEG and eye-tracking data (Pickle format), please navigate to:\n  `derivatives/synced/`\n* **Linguistic Content:** Reading materials (words with coordinate information) are stored in `derivatives/stimuli/wiki_stories`. Each word is assigned a unique key to enable mapping fixated words back to their original corpus.\n* **Raw EEG (BIDS):** **Work in Progress.** We are currently converting the full raw EEG dataset for all participants into BIDS-compliant format. \n\n## Project Details\n- **Task:** Naturalistic reading of standardized articles with retrospective self-report paradigm (ReMind task).\n- **Participants:** 44 subjects (50+ hours of data).\n- **Modalities:** \n    - EEG (BioSemi ActiveTwo 64 channels).\n    - Simultaneous Eye-Tracking (SR Research EyeLink 1000 Plus).\n    - Span-level mind-wandering annotations. \n    - Reading comprehension scores (page-level, multiple-choice questions).\n\n## Structure\nThis repository follows the Brain Imaging Data Structure (BIDS). \n- `participants.tsv`: Demographic information (age, sex, handedness, ADHD/Reading Disability status).\n- `derivatives/synced/`: Synchronized multi-modal data frames ready for Machine Learning pipelines.\n\n## Publication & Citation\nThe dataset paper describing the collection, synchronization, and baseline modeling of this data will be available online shortly. Once published, please use the citation provided here to credit the work.\n\n\n\n\n","bids_version":"1.9.0","sessions_count":null,"publish_date":null,"embedding_dirty":0,"license_tier":"public","zarr_status":"ready","zarr_converted_at":"2026-08-23 02:47:20","zarr_store_count":5,"zarr_index_etag":"2f7f4a3bbdf20bd3e6d0a2fb0677a495","zarr_source_commit":"43f5de7f91b2258b656fc8a15d47ffbe6efc6832","archive_status":"ready","archive_size":21949330557,"archive_retry_count":0,"records_status":"ready","archive_skip_reason":null,"zarr_errors":0,"zarr_failure_count":0,"zarr_deterministic":0,"zarr_failed_at":null,"num_dataset_citations":2,"num_datapaper_citations":0,"n_channels":64,"electrode_system":"10-10","has_hed":0,"hed_version":null,"is_exemplar":0,"bytes_present":50023448739,"data_complete":1,"withdrawn_at":null,"withdrawn_reason":null,"archive_complete":null,"archive_absent_files":null,"archive_declared_files":null,"zarr_pool_breaks":null,"total_recording_duration":3551,"recording_duration_min":659,"recording_duration_max":814,"recording_count":5,"recordings_unavailable":0,"recordings_measured":5,"channel_count_min":64,"channel_count_max":64,"sampling_frequency":256,"power_line_frequency":60,"eeg_reference":"CMS/DRL","placement_scheme":"BioSemi 64 system","sweep_stamps":"{\"enrichment_updated_at\":\"2026-08-18 23:39:27\",\"metadata_updated_at\":\"2026-08-18 23:39:34\",\"archive_checked_at\":\"2026-06-30 16:29:45\",\"zarr_checked_at\":null,\"records_checked_at\":\"2026-06-30 16:11:26\",\"citations_updated_at\":\"2026-09-08 03:00:52\",\"channel_montage_checked_at\":null,\"hed_checked_at\":null,\"data_checked_at\":\"2026-09-04 03:00:28\",\"availability_report_at\":\"2026-07-23 01:33:12\",\"recording_stats_at\":\"2026-09-02 11:34:04\",\"signal_defaults_at\":\"2026-09-02 12:56:40\"}","participants":1,"num_citations":2,"latest_version":"v1.0.0","zarr_verify_status":null,"zarr_verified_at":null,"owner_username":"nemarAdmin","owner_github":"nemarAdmin","file_size_formatted":"46.59 GB","zarr_data_failures":null,"zarr_index_url":"https://zarr.nemar.org/on007629/zarr/index.json","attestation_deposit_type":null,"attestation_key_status":null,"attestation_deidentified":null,"attestation_no_duplicate":null,"attestation_upstream_source":null,"attestation_accepted_at":null}}