{"dataset":{"id":"61269","dataset_id":"on007753","name":"BCCWJ-EEG","description":"This dataset contains 64-channel EEG recordings from 41 Japanese native speakers who read Japanese newspaper articles from the Balanced Corpus of Contemporary Written Japanese (BCCWJ), presented word by word via rapid serial visual presentation. It is part of the BCCWJ-Brain collection, which also includes fMRI and MEG datasets acquired from separate participant groups using the same stimuli, enabling cross-modal comparisons of language processing with complementary spatial and temporal resolution.","owner_user_id":15,"status":"active","github_repo":"nemarDatasets/on007753","concept_doi":"10.82901/nemar.on007753","latest_version_doi":"10.82901/nemar.on007753.v1.0.0","created_at":"2026-07-06 02:01:28","updated_at":"2026-08-18 22:39:24","zenodo_concept_id":null,"is_sandbox":0,"visibility":"public","ezid_status":"public","enrichment_json":"{\n  \"version\": \"2.0\",\n  \"pipeline_stage\": \"validated\",\n  \"title\": \"BCCWJ-EEG\",\n  \"description\": \"This dataset contains 64-channel EEG recordings from 41 Japanese native speakers who read Japanese newspaper articles from the Balanced Corpus of Contemporary Written Japanese (BCCWJ), presented word by word via rapid serial visual presentation. It is part of the BCCWJ-Brain collection, which also includes fMRI and MEG datasets acquired from separate participant groups using the same stimuli, enabling cross-modal comparisons of language processing with complementary spatial and temporal resolution.\",\n  \"methods_description\": \"EEG data were recorded using a BrainAmp amplifier with a 64-channel electrode cap, referenced online at FCz with ground at AFz, and an infraorbital electrode for ocular artifact monitoring. Impedances were kept below 20kΩ and data were sampled at 1,000Hz. Stimuli (20 newspaper articles) were presented word by word via RSVP in PsychoPy, with 500 ms presentation and 500 ms blank intervals. Preprocessing was performed using MNE-Python and Eelbrain, including ICA-based ocular artifact removal, 0.1–40 Hz bandpass filtering, re-referencing to common average, epoching from -100 to 1000 ms relative to word onset, downsampling to 200 Hz, and baseline correction.\",\n  \"license\": \"CC0\",\n  \"dataset_type\": \"raw\",\n  \"authors\": {\n    \"Yushi Sugimoto\": {},\n    \"Masayuki Asahara\": {},\n    \"Hyeonjeong Jeong\": {},\n    \"Akitake Kanno\": {},\n    \"Masatoshi Koizumi\": {},\n    \"Yohei Oseki\": {}\n  },\n  \"keywords\": [\n    {\n      \"term\": \"EEG\"\n    },\n    {\n      \"term\": \"Natural Language Processing\",\n      \"subject_scheme\": \"MeSH\",\n      \"scheme_uri\": \"https://id.nlm.nih.gov/mesh/\",\n      \"value_uri\": \"http://id.nlm.nih.gov/mesh/D009323\"\n    },\n    {\n      \"term\": \"Reading\",\n      \"subject_scheme\": \"MeSH\",\n      \"value_uri\": \"http://id.nlm.nih.gov/mesh/D011932\"\n    },\n    {\n      \"term\": \"Japanese\"\n    },\n    {\n      \"term\": \"psycholinguistics\"\n    },\n    {\n      \"term\": \"rapid serial visual presentation\"\n    }\n  ],\n  \"related_identifiers\": [\n    {\n      \"identifier\": \"https://github.com/nemarDatasets/on007753\",\n      \"identifier_type\": \"URL\",\n      \"relation_type\": \"IsDescribedBy\"\n    },\n    {\n      \"identifier\": \"https://nemar.org/dataset/on007753\",\n      \"identifier_type\": \"URL\",\n      \"relation_type\": \"IsDescribedBy\"\n    },\n    {\n      \"identifier\": \"10.7554/eLife.85012\",\n      \"identifier_type\": \"DOI\",\n      \"relation_type\": \"References\"\n    },\n    {\n      \"identifier\": \"10.3389/fnins.2013.00267\",\n      \"identifier_type\": \"DOI\",\n      \"relation_type\": \"References\"\n    },\n    {\n      \"identifier\": \"10.1007/s10579-013-9261-0\",\n      \"identifier_type\": \"DOI\",\n      \"relation_type\": \"References\"\n    },\n    {\n      \"identifier\": \"10.18112/openneuro.ds007753.v1.1.1\",\n      \"identifier_type\": \"DOI\",\n      \"relation_type\": \"IsDerivedFrom\"\n    },\n    {\n      \"identifier\": \"10.1016/j.jneumeth.2006.11.017\",\n      \"identifier_type\": \"DOI\",\n      \"relation_type\": \"References\"\n    },\n    {\n      \"identifier\": \"10.3389/neuro.11.010.2008\",\n      \"identifier_type\": \"DOI\",\n      \"relation_type\": \"References\"\n    }\n  ],\n  \"resource_type_specific\": \"EEG Dataset\",\n  \"modalities\": [\n    \"eeg\"\n  ],\n  \"sizes\": [\n    \"94.6 GB (290 files)\"\n  ],\n  \"formats\": [\n    \".eeg\",\n    \".fif\",\n    \".json\",\n    \".md\",\n    \".tsv\",\n    \".vhdr\",\n    \".vmrk\",\n    \".yml\"\n  ],\n  \"source_hash\": \"9a796ad4631d6b786bb43a48f1a7bb3ca30e1c802dab4ec5ce5fcab0e31717c1\"\n}","last_activity_at":"2026-07-06 02:01:28","source":"openneuro","source_id":"ds007753","subject_count":41,"modalities":"eeg","age_min":18,"age_max":35,"file_size":95701291014,"total_files":505,"tasks":"BCCWJreading","metadata_columns_error":null,"staleness_warn_stage":null,"staleness_admin_notified_at":null,"authors":"Yushi Sugimoto, Masayuki Asahara, Hyeonjeong Jeong, Akitake Kanno, Masatoshi Koizumi, Yohei Oseki","license":"CC0","readme":"[![DOI](https://img.shields.io/badge/DOI-10.82901%2Fnemar.on007753-blue)](https://doi.org/10.82901/nemar.on007753)\n\n\n## Overview\nThis dataset contains EEG data collected while Japanese native speakers read Japanese newspaper articles from the Balanced Corpus of Contemporary Written Japanese (BCCWJ; Maekawa et al., 2014). Stimuli were presented word by word. This dataset is part of BCCWJ-Brain, This dataset is part of BCCWJ-Brain; three types of brain data (fMRI, MEG, and EEG) were acquired from separate groups of participants using the same stimuli, enabling cross-modality comparisons of language processing with high spatial and temporal resolution respectively.\nThe BCCWJ-Brain collection consists of the following datasets:\n\n- BCCWJ-fMRI: ds007752 ([https://openneuro.org/datasets/ds007752](https://openneuro.org/datasets/ds007752))\n- BCCWJ-MEG: ds007763 ([https://openneuro.org/datasets/ds007763](https://openneuro.org/datasets/ds007763))\n- BCCWJ-EEG: ds007753 ([https://openneuro.org/datasets/ds007753](https://openneuro.org/datasets/ds007753))\n\n\n## Participants\nData from forty-one participants were included in the dataset (22 females and 19 males; mean age=20.5 (SD = 2.89)). All participants were right-handed, had no neurological illness, and had normal or corrected-to-normal vision. \n\n## Data Acquisition\nEEG data were recorded using a BrainAmp amplifier (Brain Products GmbH, Germany) with a 64-channel electrode cap. The online reference electrode was placed at FCz, and the ground electrode was placed at AFz. An electrode was placed below the right eye (IO) to monitor ocular artifacts. Electrode impedances were kept below 20kΩ prior to the recording. Data were recorded at a sampling rate of 1,000Hz.\n\n## Experiment Procedure\nTwenty Japanese newspaper articles were used as stimuli. Stimuli were presented word by word using rapid serial visual presentation (RSVP) implemented in PsychoPy (Peirce, 2007, 2009). Each stimulus was presented for 500 ms, followed by a 500 ms blank screen. The order of the newspaper articles were randomized.\n\n## Data Preprocessing\nEEG data were preprocessed using MNE-Python (v1.9.0; Gramfort et al., 2013) and Eelbrain (v0.40.3; Brodbeck et al., 2023). Continuous EEG was recorded with an easycap-M1 electrode layout and one bipolar EOG channel (IO). Independent component analysis (ICA) was then applied, and components reflecting ocular artifacts were identified and removed. The ICA decomposition was subsequently applied to the 0.1–40 Hz bandpass filtered data, and the cleaned signal was re-referenced to the common average. Data were then segmented into epochs from −100 to 1,000 ms relative to word onset and downsampled to 200 Hz. Baseline correction was applied using the pre-stimulus interval (−100 to 0 ms). The preprocessed EEG data can be found in the ``derivatives``.\n\n```\nderivatives/\n├── eeg/\n    ├── sub-XX_task-BCCWJreading_eeg_0.1-40_raw.fif (bandpass filter applied files)\n    ├── sub-XX_task-BCCWJreading_eeg_raw.fif (raw data converted to fif files (without downsampling (1,000Hz)))\n    ├── sub-XX_task-BCCWJreading_eeg_reref0.1-40-ica_raw.fif (preprocessed files)\n    └── sub-XX_task-BCCWJreading_eeg_reref0.1-40-ica_ave.fif (evoked files)\n```\n\n## Notes\nSince the BCCWJ texts are not copyright-free, texts for the experiment is not included in this dataset. To obtain the text, users must register for access to BCCWJ ([https://bccwj-data.ninjal.ac.jp/](https://bccwj-data.ninjal.ac.jp/)) separately.  See [https://clrd.ninjal.ac.jp/bccwj/en/subscription.html](https://clrd.ninjal.ac.jp/bccwj/en/subscription.html) for the details. Once access is granted, we provide a script to incorporate the text into the corresponding ``events.tsv`` files.\n\n\n## References\nBrodbeck, C., Das, P., Gillis, M., Kulasingham, J. P., Bhattasali, S., Gaston, P., Resnik, P., & Simon, J. Z. (2023). Eelbrain, a Python toolkit for time-continuous analysis with temporal response functions. eLife, 12, e85012. https://doi.org/10.7554/eLife.85012\n\nGramfort, A., Luessi, M., Larson, E., Engemann, D. A., Strohmeier, D., Brodbeck, C., Goj, R., Jas, M., Brooks, T., Parkkonen, L., & Hämäläinen, M. (2013). MEG and EEG data analysis with MNE-Python. Frontiers in Neuroinformatics, 7, 267. https://doi.org/10.3389/fnins.2013.00267\n\nMaekawa, K., Yamazaki, M., Ogiso, T., Maruyama, T., Ogura, H., Kashino, W., Koiso, H., Yamaguchi, M., Tanaka, M., & Den, Y. (2014). Balanced corpus of contemporary written Japanese. Language Resources and Evaluation, 48, 345–371. https://doi.org/10.1007/s10579-013-9261-0\n\nPeirce, J. W. (2007). PsychoPy—Psychophysics software in Python. Journal of Neuroscience Methods, 162(1–2), 8–13. https://doi.org/10.1016/j.jneumeth.2006.11.017\n\nPeirce, J. W. (2009). Generating stimuli for neuroscience using PsychoPy. Frontiers in Neuroinformatics, 2, 10. https://doi.org/10.3389/neuro.11.010.2008","bids_version":"1.9.0","sessions_count":null,"publish_date":null,"embedding_dirty":0,"license_tier":"public","zarr_status":"ready","zarr_converted_at":"2026-08-27 02:55:07","zarr_store_count":41,"zarr_index_etag":"a07e2d6fc776f489f8a7a43139e3d150","zarr_source_commit":"d6c5e160408a21663033984c0188136cd66e18a4","archive_status":"ready","archive_size":70219074463,"archive_retry_count":0,"records_status":"ready","archive_skip_reason":null,"zarr_errors":0,"zarr_failure_count":0,"zarr_deterministic":0,"zarr_failed_at":null,"num_dataset_citations":1,"num_datapaper_citations":0,"n_channels":64,"electrode_system":"10-10","has_hed":0,"hed_version":null,"is_exemplar":0,"bytes_present":94588307892,"data_complete":1,"withdrawn_at":null,"withdrawn_reason":null,"archive_complete":null,"archive_absent_files":null,"archive_declared_files":null,"zarr_pool_breaks":0,"total_recording_duration":91215.31999999998,"recording_duration_min":2123.18,"recording_duration_max":2468.76,"recording_count":41,"recordings_unavailable":0,"recordings_measured":41,"channel_count_min":64,"channel_count_max":64,"sampling_frequency":1000,"power_line_frequency":50,"eeg_reference":null,"placement_scheme":null,"sweep_stamps":"{\"enrichment_updated_at\":\"2026-08-18 22:39:09\",\"metadata_updated_at\":\"2026-08-18 22:39:22\",\"archive_checked_at\":\"2026-07-06 02:57:44\",\"zarr_checked_at\":null,\"records_checked_at\":\"2026-07-06 02:11:39\",\"citations_updated_at\":\"2026-09-10 03:00:21\",\"channel_montage_checked_at\":null,\"hed_checked_at\":null,\"data_checked_at\":\"2026-09-05 03:00:33\",\"availability_report_at\":\"2026-07-23 01:33:42\",\"recording_stats_at\":\"2026-09-02 11:34:07\",\"signal_defaults_at\":\"2026-09-02 12:58:14\"}","participants":41,"num_citations":1,"latest_version":"v1.0.0","zarr_verify_status":null,"zarr_verified_at":null,"owner_username":"nemarAdmin","owner_github":"nemarAdmin","file_size_formatted":"89.13 GB","zarr_data_failures":null,"zarr_index_url":"https://zarr.nemar.org/on007753/zarr/index.json","attestation_deposit_type":null,"attestation_key_status":null,"attestation_deidentified":null,"attestation_no_duplicate":null,"attestation_upstream_source":null,"attestation_accepted_at":null}}