Please use this identifier to cite or link to this item: http://nopr.niscpr.res.in/handle/123456789/68236
metadata.dc.identifier.doi: https://doi.org/10.56042/jsir.v85i4.22732
Title: Seeing Beyond Text: A Visual-Linguistic Dataset and Multimodal Framework for English-Hindi Video-Guided Translation
Authors: Paul, Binnu
Rudrapal, Dwijen
Chakma, Kunal
Jamatia, Anupam
Keywords: Artificial intelligence;Cross-lingual learning;Indian languages;Machine learning;Natural language processing
Issue Date: Apr-2026
Publisher: NIScPR-CSIR, India
Abstract: Despite the progress in Neural Machine Translation (NMT), translating ambiguous and context rich content remains a major challenge, especially in low-resource language pairs like English-Hindi. Traditional NMT systems often fail to solve these challenges due to their reliance on textual data alone. Multimodal approaches, particularly those incorporating visual context, offer promising solutions to the task by resolving linguistic ambiguities. To address this, a novel solution is introduced through a Visual Scene-Aware Hindi Subtitles Dataset (VISA-HIN), designed specifically for English-Hindi Video-Guided Multimodal Machine Translation (VMMT). This dataset aligns English subtitles with corresponding video frames and provides Hindi translations. Alongside the dataset, this study propose a video-guided MMT framework that leverages visual cues to enhance translation quality. The results of the experiments show the potential of scene aware information to improve contextual understanding and fluency in English-to-Hindi translation, paving the way for more robust and accurate multimodal translation systems in low-resource settings.
Page(s): 312-324
ISSN: 0975-1084 (Online) ; 0022-4456 (Print)
Appears in Collections:JSIR Vol.85(04) [April 2026]

Files in This Item:
File Description SizeFormat 
JSIR 85(4) 312-324.pdf1.49 MBAdobe PDFView/Open


Items in NOPR are protected by copyright, with all rights reserved, unless otherwise indicated.