EGOEXTRA Dataset - Release Documentation This directory contains the released version of the EGOEXTRA dataset, organized into video clips, transcripts, and question–answer annotations across multiple scenarios. ROOT STRUCTURE Dataset_release/ |-- video_clip_final/ |-- transcript/ |-- final_questions/ Each subfolder is described in detail below. --- final_questions/ This folder contains the final question–answer annotations, organized as one file per scenario. Example: final_questions/ |-- 01_keep_with_timestamp_filtered.json |-- 02_review_with_timestamp_filtered.json |-- ... Each JSON file corresponds to a specific scenario (e.g., 01 = KEEP) and contains a list of questions. QUESTION JSON FORMAT Each question entry follows this structure: { "conversation_turn": [], "context_previous": [{}], "context_after": [], "generated_question": "What object is in front of me?", "options": [], "correct_answer": "An orange bicycle", "questions_involved_id": [], "univoc_id": "01_001_1", "question_start_time": "00:00:02,000", "question_end_time": "00:00:12,480" } FIELD DESCRIPTION conversation_turn The conversation turn on which the question is primarily based. context_previous Conversation turns immediately preceding the main turn. context_after Conversation turns immediately following the main turn. generated_question The natural language question generated. options A list of five multiple-choice options. correct_answer The correct answer among the provided options. questions_involved_id IDs of transcript utterances (lines) involved in the question and answer. univoc_id A unique identifier for the question (scenario-specific). question_start_time / question_end_time Start and end timestamps of the question with respect to the raw video. --- transcript/ This folder contains the raw transcripts associated with the videos. * One transcript file per video * Includes the original conversational turns * Transcript IDs are referenced by questions_involved_id in the question annotations Example transcript entries: ID 17 [00:00:25,980 --> 00:00:31,690] E: You have to lift the bicycle and the tube holding the saddle you have to insert it between the grip ID 18 [00:00:28,280 --> 00:00:36,660] A: Yes? So in this way? The timestamps [00:00:25,980 --> 00:00:31,690] indicate the start and end time in the raw video. --- video_clip_final/ This folder contains the video clips associated with each question. Key characteristics: * Clips are extracted before the question timestamp * Multiple temporal resolutions are provided * Typical clip durations include: * 5 seconds * 15 seconds * 30 seconds These clips are designed to support temporal reasoning and varying context lengths for video-based question answering. Since 2 or more question can be close in time, we used a single video clip. This association is described in clips_mapping__duration_<5,15,or 30> --- NOTES * All timestamps are aligned with the raw video timeline. * Scenario numbering is consistent across videos, transcripts, and question files. * The dataset supports research in egocentric video understanding, question answering, and temporal reasoning. --- Commercial Use Restriction This dataset may not be used, in whole or in part, for any commercial purpose, including but not limited to training models for commercial products, paid services, or monetized applications, without prior written consent from the dataset authors. For any questions michele.mazzamuto@phd.unict.it