{"dataset":{"id":"149","dataset_id":"nm000133","name":"Alljoined1","description":null,"owner_user_id":15,"status":"active","github_repo":"nemarDatasets/nm000133","concept_doi":"10.82901/nemar.nm000133","latest_version_doi":"10.82901/nemar.nm000133.v1.0.2","created_at":"2026-03-15 07:33:10","updated_at":"2026-07-10 21:46:23","zenodo_concept_id":"20275887","is_sandbox":0,"visibility":"public","ezid_status":"public","enrichment_json":"{\n  \"version\": \"2.0\",\n  \"pipeline_stage\": \"enriched\",\n  \"title\": \"Alljoined1\",\n  \"description\": \"Alljoined1 is a 64-channel EEG dataset comprising neural responses from eight healthy adults (6 male, 2 female; mean age 22 ± 0.64 years) viewing 10,000 natural images each during rapid serial visual presentation (RSVP) with an oddball detection task. The dataset contains approximately 46,080 epochs across 13 sessions, with stimuli drawn from the Natural Scenes Dataset and MS-COCO. Raw data are preserved in 24-bit BioSemi Data Format, and preprocessed epoched derivatives are provided in MNE-Python format, enabling research in EEG-to-image decoding and visual neuroscience.\",\n  \"methods_description\": \"EEG recordings were acquired using a 64-channel BioSemi ActiveTwo system (Ag/AgCl sintered electrodes, International 10-20 montage) at 512 Hz sampling rate with 24-bit A/D conversion. Participants viewed natural images in an RSVP paradigm (300 ms image presentation, 300 ms black screen, 0-50 ms jitter) while performing an oddball detection task. Preprocessing included 0.5-125 Hz band-pass filtering, 60 Hz notch filtering, FastICA (95% variance retained), epoching (-50 to 600 ms), AutoReject artifact rejection, baseline correction, and average re-referencing.\",\n  \"license\": \"CC-BY-NC-ND-4.0\",\n  \"dataset_type\": \"raw\",\n  \"authors\": {\n    \"Jonathan Xu\": {},\n    \"Si Kai Lee\": {},\n    \"Wangshu Jiang\": {}\n  },\n  \"keywords\": [\n    {\n      \"term\": \"EEG\"\n    },\n    {\n      \"term\": \"visual perception\"\n    },\n    {\n      \"term\": \"image decoding\"\n    },\n    {\n      \"term\": \"rapid serial visual presentation\"\n    },\n    {\n      \"term\": \"natural scenes\"\n    },\n    {\n      \"term\": \"Natural Scenes Dataset\"\n    },\n    {\n      \"term\": \"MS-COCO\"\n    },\n    {\n      \"term\": \"brain-computer interfaces\"\n    },\n    {\n      \"term\": \"oddball detection\"\n    },\n    {\n      \"term\": \"BioSemi\"\n    },\n    {\n      \"term\": \"artifact rejection\"\n    },\n    {\n      \"term\": \"independent component analysis\"\n    },\n    {\n      \"term\": \"preprocessing\"\n    },\n    {\n      \"term\": \"visual neuroscience\"\n    },\n    {\n      \"term\": \"neural decoding\"\n    }\n  ],\n  \"related_identifiers\": [\n    {\n      \"identifier\": \"10.48550/arXiv.2404.05553\",\n      \"identifier_type\": \"DOI\",\n      \"relation_type\": \"References\"\n    },\n    {\n      \"identifier\": \"https://github.com/nemarDatasets/nm000133\",\n      \"identifier_type\": \"URL\",\n      \"relation_type\": \"IsDescribedBy\"\n    },\n    {\n      \"identifier\": \"https://nemar.org/dataexplorer/detail?dataset_id=nm000133\",\n      \"identifier_type\": \"URL\",\n      \"relation_type\": \"IsDescribedBy\"\n    },\n    {\n      \"identifier\": \"10.82901/nemar.nm000133\",\n      \"identifier_type\": \"DOI\",\n      \"relation_type\": \"IsVersionOf\"\n    }\n  ],\n  \"resource_type_general\": \"Dataset\",\n  \"resource_type_specific\": \"EEG Dataset\",\n  \"modalities\": [\n    \"eeg\"\n  ],\n  \"sizes\": [\n    \"8.3 GB (1002 files)\"\n  ],\n  \"formats\": [\n    \".DS_Store\",\n    \".bdf\",\n    \".csv\",\n    \".fif\",\n    \".gitignore\",\n    \".ipynb\",\n    \".jpg\",\n    \".json\",\n    \".m\",\n    \".mat\",\n    \".md\",\n    \".png\",\n    \".py\",\n    \".sh\",\n    \".tsv\",\n    \".yml\"\n  ],\n  \"source_hash\": \"beb18cbcc5d7d5c490cac3e8c506fd4a2a507fc0aee0b21886dfc355ccbced5e\"\n}","last_activity_at":"2026-05-01 18:28:17","source":null,"source_id":null,"subject_count":8,"modalities":"eeg","age_min":null,"age_max":null,"file_size":8315857513,"total_files":1002,"tasks":"images","metadata_columns_error":null,"staleness_warn_stage":null,"staleness_admin_notified_at":null,"authors":"Jonathan Xu, Si Kai Lee, Wangshu Jiang","license":"CC-BY-NC-ND-4.0","readme":"[![DOI](https://img.shields.io/badge/DOI-10.82901%2Fnemar.nm000133-blue)](https://doi.org/10.82901/nemar.nm000133)\n\n# Alljoined1: EEG Responses to Natural Images\n\n## Overview\n\nAlljoined1 is an EEG dataset of neural responses to rapid serial visual presentation (RSVP) of natural images, designed for EEG-to-image decoding research. Eight healthy right-handed adults (6 male, 2 female; mean age 22 +/- 0.64 years, normal or corrected-to-normal vision) each viewed 10,000 natural images across two recording sessions on separate days.\n\nThe original data were recorded in BioSemi Data Format (BDF) via a 64-channel BioSemi ActiveTwo system with 24-bit A/D conversion, digitized at 512 Hz. This BIDS-formatted version preserves the BDF format to maintain full 24-bit data fidelity.\n\n**Reference:** Xu, J., Lee, S. K., & Jiang, W. (2024). Alljoined -- A dataset for EEG-to-Image decoding. <https://doi.org/10.48550/arXiv.2404.05553>\n\n## Recording Setup\n\n- **Equipment:** BioSemi ActiveTwo, 64 Ag/AgCl sintered electrodes\n- **Montage:** International 10-20 system\n- **Sampling rate:** 512 Hz\n- **Reference:** CMS/DRL (BioSemi default); average reference applied in preprocessing\n- **Electrode offset:** kept below 40 mV\n- **Power line:** 60 Hz notch filter applied during preprocessing\n\n## Task Paradigm\n\nParticipants viewed natural images in a rapid serial visual presentation (RSVP) paradigm with an oddball detection task. Each trial consisted of an image presented for 300 ms, followed by 300 ms of black screen, plus 0-50 ms of random jitter. Participants pressed the space bar when two consecutive trials contained the same image (oddball detection). Oddball trials (24 per block) were excluded from analysis.\n\n## Stimulus Set\n\n10,000 natural images per participant drawn from the Natural Scenes Dataset (NSD), which itself is sourced from MS-COCO:\n\n- **1,000 shared images:** the first 960 images from the NSD \"shared1000\" subset, shown to all participants (each image repeated 4 times per participant)\n- **9,000 unique images:** different for each participant\n\nEach image was shown 4 times per participant across blocks and sessions (presented twice per block, with blocks repeated within sessions).\n\nThe BIDS event tables (every `events.tsv`) reference the stimuli as\n`trial_type = \"image/N\"` where `N` is a 1-indexed position (1..960) into\nthe shared subset. The full mapping chain is:\n\n```\nevents.tsv `value` (N, 1..960)\n    ↓\nsharedix[N-1]                                    (from code/0_data_collection/nsd_expdesign.mat; 1-indexed NSD id)\n    ↓\nnsdId = sharedix[N-1] - 1                        (0-indexed)\n    ↓\ncode/1_preprocessing/data/nsd_stim_info_merged.csv\n    ↓\ncocoId, cocoSplit (val2017 / train2017)\n    ↓\nstimuli/<cocoSplit>/000000<cocoId:012d>.jpg      (preserves original COCO 2017 layout)\n```\n\nEmpirically the 960 shared NSD ids are all in `train2017`, so every\nstimulus path under this dataset is `stimuli/train2017/000000<id>.jpg`.\n\nTo populate `stimuli/`:\n\n```bash\npython code/download_stimuli.py            # ~140 MB, fetches only the 960 needed\npython code/smoke_test.py                  # confirms every event row resolves\n```\n\nA small alignment helper is provided:\n\n```python\nimport pandas as pd\nfrom code.align_stimuli import StimulusAligner\n\naligner = StimulusAligner('.')\nevents = pd.read_csv('sub-01/ses-01/eeg/sub-01_ses-01_task-images_events.tsv', sep='\\t')\npaths = aligner.paths_for_events(events, subject=1, session=1)   # list[Path | None]\nimg   = aligner.image_for_event(events.iloc[0], subject=1, session=1)  # PIL.Image\n```\n\n## Subjects and Sessions\n\n8 subjects, 1-2 sessions each (13 sessions total):\n\n| Subject | Sessions | Notes |\n|---------|----------|-------|\n| sub-01 | ses-01, ses-02 | |\n| sub-02 | ses-01 | |\n| sub-03 | ses-01, ses-02 | Epoched file missing for ses-01 |\n| sub-04 | ses-01, ses-02 | |\n| sub-05 | ses-01, ses-02 | |\n| sub-06 | ses-01, ses-02 | |\n| sub-07 | ses-01 | |\n| sub-08 | ses-01 | |\n\nTotal: approximately 46,080 epochs across all participants (approximately 3,839 events per session after oddball exclusion).\n\n## Data Format\n\nRaw continuous EEG recordings are stored as BDF files (BioSemi Data Format, 24-bit resolution). The original data were distributed as MNE-Python FIF files; conversion to BDF was performed to preserve the native 24-bit precision of the BioSemi ActiveTwo system. Round-trip validation confirmed data integrity to within 1.55e-8 V (sub-nanovolt), and event onsets match exactly (zero timing error).\n\n**Per-session files:**\n\n| Path | Description |\n|------|-------------|\n| `sub-XX/ses-YY/eeg/sub-XX_ses-YY_task-images_eeg.bdf` | Raw EEG |\n| `sub-XX/ses-YY/eeg/sub-XX_ses-YY_task-images_events.tsv` | Event markers |\n\n**Shared sidecar files (root level, BIDS inheritance principle):**\n\n| File | Description |\n|------|-------------|\n| `task-images_eeg.json` | Recording parameters |\n| `task-images_channels.tsv` | Channel descriptions (64 EEG channels) |\n| `task-images_electrodes.tsv` | Electrode positions (standard 10-20, CapTrak) |\n| `task-images_coordsystem.json` | Coordinate system specification |\n\nEvent values in the events.tsv files represent image indices (1-960+) corresponding to NSD image identifiers. The `trial_type` column uses the format `image/{index}`.\n\n## Derivatives\n\nThe `derivatives/epoched/` directory contains preprocessed and epoched data provided by the original authors, stored in MNE-Python FIF format (`.fif`).\n\nPreprocessing pipeline applied by the original authors:\n\n1. Band-pass filter: 0.5-125 Hz\n2. Notch filter: 60 Hz (power line)\n3. Independent Component Analysis (ICA): FastICA, retaining 95% of variance\n4. Epoch extraction: -50 ms to 600 ms relative to stimulus onset\n5. Artifact rejection: AutoReject algorithm (mean 130.75 epochs dropped per subject, SD 260.44)\n6. Baseline correction\n7. Average re-referencing\n\nThese epoched files are derivative products, not raw recordings, and are stored separately per BIDS conventions. Note: the epoched file for sub-03 ses-01 was not available in the source distribution.\n\n## Code\n\nThe `code/` directory contains the original Alljoined1 analysis code, cloned from <https://github.com/Alljoined/alljoined-dataset1>.\n\n## BIDS Conversion\n\nConverted to BIDS by Yahya Shirazi (Swartz Center for Computational Neuroscience, UC San Diego) using MNE-Python and custom scripts.\n\n- **Source data:** OSF repository <https://osf.io/kqgs8/>\n- Conversion validated with round-trip integrity checks (data, channels, sampling frequency, event count, event values, and event timing)\n\n## License and Terms of Use\n\nThis dataset is distributed under CC-BY-NC-ND-4.0 (Creative Commons Attribution-NonCommercial-NoDerivatives 4.0). The Alljoined team imposes additional terms on their datasets. By using this dataset you agree to all conditions below.\n\n1. Researcher shall use the Dataset only for non-commercial research and educational purposes, in accordance with Alljoined's [Terms of Use](https://www.alljoined.com/terms-of-use).\n2. **No Warranties:** Alljoined makes no representations or warranties regarding the Dataset, including but not limited to warranties of non-infringement or fitness for a particular purpose.\n3. **Full Responsibility:** Researcher accepts full responsibility for his or her use of the Dataset and shall defend and indemnify Alljoined, including their employees, officers and agents, against any and all claims arising from Researcher's use of the Dataset.\n4. **Privacy Compliance:** Researcher shall comply with Alljoined's [Privacy Policy](https://www.alljoined.com/privacy-policy) and ensure that any use of the Dataset respects the privacy rights of individuals whose data may be included.\n5. **Sharing Rights:** Researcher may provide research associates and colleagues with access to the Dataset provided that they first agree to be bound by these terms and conditions.\n6. **Termination Rights:** Alljoined reserves the right to terminate Researcher's access to the Dataset at any time.\n7. **Commercial Entity Binding:** If Researcher is employed by a for-profit, commercial entity, Researcher's employer shall also be bound by these terms and conditions, and Researcher hereby represents that he or sh","bids_version":"1.9.0","sessions_count":2,"publish_date":"2026-03-17 20:59:32","embedding_dirty":0,"license_tier":"noderiv","zarr_status":"ready","zarr_converted_at":"2026-09-04 13:02:38","zarr_store_count":13,"zarr_index_etag":"f00c56be6e7dd768857a8b3b5c745978","zarr_source_commit":"f87b9a276404e261c172afbae47161b3d8003f9e","archive_status":"ready","archive_size":7196232622,"archive_retry_count":0,"records_status":null,"archive_skip_reason":null,"zarr_errors":0,"zarr_failure_count":0,"zarr_deterministic":0,"zarr_failed_at":null,"num_dataset_citations":0,"num_datapaper_citations":6,"n_channels":null,"electrode_system":null,"has_hed":0,"hed_version":null,"is_exemplar":0,"bytes_present":null,"data_complete":null,"withdrawn_at":null,"withdrawn_reason":null,"archive_complete":null,"archive_absent_files":null,"archive_declared_files":null,"zarr_pool_breaks":0,"total_recording_duration":44932,"recording_duration_min":3251,"recording_duration_max":3488,"recording_count":13,"recordings_unavailable":0,"recordings_measured":13,"channel_count_min":64,"channel_count_max":64,"sampling_frequency":512,"power_line_frequency":60,"eeg_reference":"CMS/DRL","placement_scheme":"10-20","sweep_stamps":"{\"enrichment_updated_at\":\"2026-05-18 18:58:49\",\"metadata_updated_at\":\"2026-05-18 18:58:49\",\"archive_checked_at\":\"2026-06-05 01:32:49\",\"zarr_checked_at\":\"2026-06-07 17:58:19\",\"records_checked_at\":null,\"citations_updated_at\":\"2026-09-08 03:00:50\",\"channel_montage_checked_at\":\"2026-06-28 22:51:25\",\"hed_checked_at\":\"2026-06-30 04:10:44\",\"data_checked_at\":null,\"availability_report_at\":\"2026-07-23 01:05:14\",\"signal_defaults_at\":\"2026-09-02 11:37:54\",\"recording_stats_at\":\"2026-09-05 03:00:34\",\"zarr_verify_attempted_at\":\"2026-09-07 03:01:13\",\"zarr_verified_at\":\"2026-09-07 03:01:13\",\"zarr_verified_commit\":\"f87b9a276404e261c172afbae47161b3d8003f9e\",\"zarr_verify_status\":\"unverifiable\",\"zarr_verify_examples\":[],\"zarr_verify_sampled\":13.0,\"zarr_verify_checked\":0.0,\"zarr_verify_checked_channels\":0.0,\"zarr_verify_checked_duration\":0.0,\"zarr_verify_checked_rate\":0.0,\"zarr_verify_unchecked\":13.0,\"zarr_verify_mismatch_count\":0.0,\"zarr_verify_examples_truncated\":0.0}","participants":8,"num_citations":6,"latest_version":"v1.0.2","zarr_verify_status":"unverifiable","zarr_verified_at":"2026-09-07 03:01:13","owner_username":"nemarAdmin","owner_github":"nemarAdmin","file_size_formatted":"7.74 GB","zarr_data_failures":null,"zarr_index_url":"https://zarr.nemar.org/nm000133/zarr/index.json","attestation_deposit_type":null,"attestation_key_status":null,"attestation_deidentified":null,"attestation_no_duplicate":null,"attestation_upstream_source":null,"attestation_accepted_at":null}}