{"dataset":{"id":"58524","dataset_id":"on004718","name":"Le Petit Prince Hong Kong: Naturalistic fMRI and EEG dataset from older Cantonese speakers","description":"This dataset presents a multimodal collection of naturalistic fMRI, structural MRI, and EEG data from 52 healthy older Cantonese-speaking adults (over 65 years) listening to excerpts from 'The Little Prince' in Cantonese. It aims to address the underrepresentation of non-Indo-European languages and older populations in language neurobiology research, providing insights into brain-behavior relationships and healthy aging. The dataset includes rich behavioral, linguistic, and cognitive measures alongside detailed audio and text annotations.","owner_user_id":15,"status":"active","github_repo":"nemarDatasets/on004718","concept_doi":"10.82901/nemar.on004718","latest_version_doi":"10.82901/nemar.on004718.v1.0.0","created_at":"2026-06-25 09:06:33","updated_at":"2026-08-19 05:16:12","zenodo_concept_id":null,"is_sandbox":0,"visibility":"public","ezid_status":"public","enrichment_json":"{\n  \"version\": \"2.0\",\n  \"pipeline_stage\": \"validated\",\n  \"title\": \"Le Petit Prince Hong Kong: Naturalistic fMRI and EEG dataset from older Cantonese speakers\",\n  \"description\": \"This dataset presents a multimodal collection of naturalistic fMRI, structural MRI, and EEG data from 52 healthy older Cantonese-speaking adults (over 65 years) listening to excerpts from 'The Little Prince' in Cantonese. It aims to address the underrepresentation of non-Indo-European languages and older populations in language neurobiology research, providing insights into brain-behavior relationships and healthy aging. The dataset includes rich behavioral, linguistic, and cognitive measures alongside detailed audio and text annotations.\",\n  \"methods_description\": \"Data were collected from 52 right-handed older Cantonese participants during separate fMRI and EEG sessions, counterbalanced in order with at least a two-week interval. In the fMRI session, participants listened to four sections of an audiobook while undergoing task-based and resting-state scanning, preceded by structural T1-weighted imaging. The EEG session similarly involved listening to four audiobook sections with comprehension questions after each run. Cognitive tasks (digit span, picture naming, verbal fluency, Flanker task) and questionnaires (LHQ3, LEQ) were also administered to assess participants' cognitive and linguistic profiles.\",\n  \"license\": \"CC0\",\n  \"dataset_type\": \"raw\",\n  \"authors\": {\n    \"Mohammad Momenian\": {},\n    \"Zhengwu Ma\": {},\n    \"Shuyi Wu\": {},\n    \"Chengcheng Wang\": {},\n    \"Jixing Li\": {}\n  },\n  \"keywords\": [\n    {\n      \"term\": \"EEG\"\n    },\n    {\n      \"term\": \"fMRI\"\n    },\n    {\n      \"term\": \"Cantonese\"\n    },\n    {\n      \"term\": \"Aging\",\n      \"subject_scheme\": \"MeSH\",\n      \"value_uri\": \"http://id.nlm.nih.gov/mesh/D000375\"\n    },\n    {\n      \"term\": \"language comprehension\"\n    },\n    {\n      \"term\": \"naturalistic stimuli\"\n    },\n    {\n      \"term\": \"older adults\"\n    },\n    {\n      \"term\": \"audiobook\"\n    },\n    {\n      \"term\": \"narrative comprehension\"\n    }\n  ],\n  \"related_identifiers\": [\n    {\n      \"identifier\": \"https://github.com/nemarDatasets/on004718\",\n      \"identifier_type\": \"URL\",\n      \"relation_type\": \"IsDescribedBy\"\n    },\n    {\n      \"identifier\": \"10.18112/openneuro.ds004718.v1.1.2\",\n      \"identifier_type\": \"DOI\",\n      \"relation_type\": \"IsDerivedFrom\"\n    },\n    {\n      \"identifier\": \"https://nemar.org/dataset/on004718\",\n      \"identifier_type\": \"URL\",\n      \"relation_type\": \"IsDescribedBy\"\n    }\n  ],\n  \"resource_type_specific\": \"Structural MRI Dataset\",\n  \"modalities\": [\n    \"anat\",\n    \"eeg\",\n    \"func\"\n  ],\n  \"sizes\": [\n    \"117.0 GB (1372 files)\"\n  ],\n  \"formats\": [\n    \".csv\",\n    \".fdt\",\n    \".gz\",\n    \".json\",\n    \".md\",\n    \".set\",\n    \".tsv\",\n    \".txt\",\n    \".wav\",\n    \".xlsx\",\n    \".yml\"\n  ],\n  \"source_hash\": \"1bf69a7e5431e6540f368486c76f2f8ce724d6f8ac4a0b7257fa4dc09fa3f9b3\"\n}","last_activity_at":"2026-06-25 09:06:33","source":"openneuro","source_id":"ds004718","subject_count":52,"modalities":"anat,eeg,func","age_min":65,"age_max":77,"file_size":117017963299,"total_files":2007,"tasks":"lppHK,rest","metadata_columns_error":null,"staleness_warn_stage":null,"staleness_admin_notified_at":null,"authors":"Mohammad Momenian, Zhengwu Ma, Shuyi Wu, Chengcheng Wang, Jixing Li","license":"CC0","readme":"[![DOI](https://img.shields.io/badge/DOI-10.82901%2Fnemar.on004718-blue)](https://doi.org/10.82901/nemar.on004718)\n\n### Update note\nSince the auditory stimuli were presented sentence by sentence, we decided to include the original audio files instead of a continuous file. We presented the story in 4 different sections. After each section, there was a time for 5 comprehension check questions. The file \"lppHK_timing_word_information.xlsx\" includes all the timing information for each section of the story. Information about each column of the file is included in the same file. We also included another file called \"EEG_trigger_and_sentence_number.xlsx\". This included information about how to match sentence ID and trigger number in the EEG data. These two files are useful for EEG data analysis. For timing issues in fMRI data analysis, we included the Eprime output files which have all the necessary information for aligning in fMRI analysis. Since Eprime usually shows some delay in the presentation of audio files, the delay could be considered in the analysis which can help with better alignment.\n\nPer OpenNeuro’s new formatting requirements, all annotation, quiz, and stimuli files are located in the “sourcedata” folder.\n\n## Overview\nIn the field of neurobiology of language, existing research predominantly focuses on data from a limited number of Indo-European languages and primarily involves younger adults, overlooking other age groups. This experiment aims to address these gaps by creating a comprehensive multimodal database. The primary goal is to advance our understanding of language processing in older adults and the impact of healthy aging on brain-behavior relationships. \n\nThe experiment involves collecting task-based and resting-state fMRI, structural MRI, and EEG data from 52 healthy right-handed older Cantonese participants over 65 years old as they listen to excerpts from “The Little Prince” in Cantonese. Additionally, the database includes detailed information on participants’ language history, lifetime experiences, linguistic and cognitive skills, as well as extensive audio and text annotations, such as time-aligned speech segmentation and prosodic features, along with word-by-word predictors from natural language processing (NLP) tools. Quality diagnostics of the MRI and EEG data confirm their robustness, positioning this database as a valuable resource for studying the spatiotemporal dynamics of language comprehension in older adults.\n\n## Methods\n### Participants\nWe recruited 52 healthy, right-handed older Cantonese participants (40 females, mean age=69.12, SD=3.52) from Hong Kong for the experiment, which consists of an fMRI and an EEG session. In both sessions, participants listened to the same sections of The Little Prince in Cantonese for approximately 20 minutes. We made sure each participant was right-handed and a native Cantonese speaker using the Language History Questionnaire8 (LHQ3). Additionally, participants reported normal or corrected normal hearing. They confirmed they had no cognitive decline. Two participants did not take part in the fMRI session and an additional 4 participants’ fMRI data were removed due to excessive head movement, resulting in a total of 46 participants (39 females, mean age=69.08yrs, SD=3.58) for the fMRI session and 52 participants (40 females, mean age=69.12yrs, SD=3.52) for the EEG session. Prior to the experiment, all participants were provided with written informed consent. All participants received monetary compensation after each session. Ethical approval was obtained from the Human Subjects Ethics Application Committee at the Hong Kong Polytechnic University (application number HSEARS20210302001). This study was performed in accordance with the Declaration of Helsinki and all other regulations set by the Ethics Committee.\n\n### Experiment Procedures\nThe study consisted of an fMRI session and an EEG session. The order of the EEG and fMRI sessions was counterbalanced across all participants, and a minimum two-week interval was maintained between sessions. \n\n#### fMRI experiment\nBefore the scanning day, an MRI safety screening form was sent to the participants to make sure MRI scanning was safe for them. We also sent them simple readings and videos about MRI scanning so that they could have an idea of what it would be like to be in a scanner. On the day of scanning, participants were initially introduced to the MRI facility and comfortably positioned inside the scanner, with their heads securely supported using paddings. An MRI-safe headphone (Sinorad package) was provided for participants to wear inside the head coil. The audio volume for the listening task was adjusted to ensure audibility for each participant. A mirror attached to the head coil allowed participants to view the stimuli presented on a screen. Participants were instructed to stay focused on the visual fixation sign while listening to the audiobook. The scanning session commenced with the acquisition of structural (T1-weighted) scans. Subsequently, participants engaged in the listening task concurrently with fMRI scanning. The task-based fMRI experiment was divided into four runs, each corresponding to a section of the audiobook. Comprehension was assessed by a series of 5 yes/no questions (20 questions in total) on the content they had listened to. These questions were presented on the screen, with participants indicating their answers by pressing a button. The session concluded with the collection of resting-state fMRI data. \n \n#### Cognitive tasks\nFour cognitive tasks were selected to assess participants’ cognitive abilities in various domains, including the forward digit span task, picture naming task, verbal fluency task, and Flanker task. These tasks were delivered after the fMRI session in a separate soundproof \nbooth.\n \n#### EEG experiment\nDuring the EEG experiment, participants were seated comfortably in a quiet room and standard procedures were followed for electrode placement and EEG cap preparation. Participants were instructed to focus on a fixation sign displayed on a monitor. The EEG recording was then initiated, with participants listening to the audiobook. The audio volume was adapted to each participant’s hearing ability before the recording using a different set of stimuli. We used Foam Ear Inserts (Medium 14mm). Similar to the fMRI experiment, participants listened to four sections of the audiobook, each lasting approximately 5 minutes. After each run, participants were asked to answer a total of 20 yes/no questions, with 5 questions assigned to each run. They indicated their answers by pressing a button. The EEG recording was conducted continuously throughout all four runs until their completion.\n\n#### Questionnaires. \nWe administered LHQ3 and the Lifetime of Experiences Questionnaire (LEQ) during EEG cap preparation. The participants did not need to move or fill in these questionnaires themselves; a research assistant asked the questions one by one in Cantonese and input the responses in an online Google form. LHQ is designed to document language history by producing aggregate scores for language proficiency, exposure, and dominance in all the languages spoken by the participants. LEQ is a tool to document what sorts of activities (e.g. sports, music, education, profession, etc) participants engage in over their lifetime. It measures lifetime experiences in three periods of life: from 13 to 30 (young adulthood), from 30 to 65 (midlife), and after 65 (late life). LEQ produces a total score (see participants.tsv) which is an indication of cognitive activity. Collecting data using these two questionnaires allowed us to have a thicker description of our participants’ linguistic, social, and cognitive experiences. \n\n### Acquisition\nThe MRI data were collected at the University Research Facility in Behavioral and Systems Neuroscience (UBSN) at The Hong Kong Polytechnic University. EEG data was collected at the Speech and Language Sciences Laboratory within the Department of Chinese and Bilingual Studies at the same university. Data acquisition for this project started in July 2021 and ended in December 2022","bids_version":"1.12.0","sessions_count":null,"publish_date":null,"embedding_dirty":0,"license_tier":"public","zarr_status":"ready","zarr_converted_at":"2026-08-11 22:20:37","zarr_store_count":96,"zarr_index_etag":"5c5144818966d8502125d3bca5e7dc8a","zarr_source_commit":"d5ec25018c5217ae1176fdae823e6eeef290ea3c","archive_status":null,"archive_size":null,"archive_retry_count":0,"records_status":"ready","archive_skip_reason":"dataset 109.0 GB exceeds 100.0 GB archive limit; use direct download","zarr_errors":6,"zarr_failure_count":6,"zarr_deterministic":1,"zarr_failed_at":"2026-08-11 22:20:37","num_dataset_citations":3,"num_datapaper_citations":0,"n_channels":64,"electrode_system":null,"has_hed":0,"hed_version":null,"is_exemplar":0,"bytes_present":117017643425,"data_complete":1,"withdrawn_at":null,"withdrawn_reason":null,"archive_complete":null,"archive_absent_files":null,"archive_declared_files":null,"zarr_pool_breaks":null,"total_recording_duration":148516.3,"recording_duration_min":1308.2,"recording_duration_max":2509.4,"recording_count":102,"recordings_unavailable":6,"recordings_measured":96,"channel_count_min":6,"channel_count_max":69,"sampling_frequency":1000,"power_line_frequency":50,"eeg_reference":"Average","placement_scheme":"10-20","sweep_stamps":"{\"enrichment_updated_at\":\"2026-08-19 05:15:54\",\"metadata_updated_at\":\"2026-08-19 05:16:10\",\"archive_checked_at\":\"2026-06-25 09:18:45\",\"zarr_checked_at\":null,\"records_checked_at\":\"2026-06-25 09:19:33\",\"citations_updated_at\":\"2026-09-08 03:00:51\",\"channel_montage_checked_at\":\"2026-06-28 23:32:36\",\"hed_checked_at\":\"2026-06-30 05:05:30\",\"data_checked_at\":\"2026-08-14 03:00:47\",\"availability_report_at\":\"2026-07-23 01:20:36\",\"recording_stats_at\":\"2026-09-02 11:32:56\",\"signal_defaults_at\":\"2026-09-02 12:19:44\"}","participants":52,"num_citations":3,"latest_version":"v1.0.0","zarr_verify_status":null,"zarr_verified_at":null,"owner_username":"nemarAdmin","owner_github":"nemarAdmin","file_size_formatted":"109 GB","zarr_data_failures":{"count":6,"detail_ref":"zarr/index.json","compacted_by":"migration_0074"},"zarr_index_url":"https://zarr.nemar.org/on004718/zarr/index.json","attestation_deposit_type":null,"attestation_key_status":null,"attestation_deidentified":null,"attestation_no_duplicate":null,"attestation_upstream_source":null,"attestation_accepted_at":null}}