Bridging Vision and Language: Advances in Image Captioning Techniques

Main Article Content

Sumedh P. Ingale ,G. R. Bamnote

Abstract

 New trends in image captioning are the central area of interest for this paper; image captioning is an area that applies computer vision and natural language processing to provide a textual explanation of an image that is descriptive semantically as well as contextually. The methods used are the Flickr 8k Dataset for obtaining high-level features using DenseNet201 trained LSTM for text generation. Some examples of preprocessing and normalization data related to texts that are important for the training and evaluation of the models are preprocessing and normalization, feature extraction, and optimization methods. Thus, the model with acceptable performance according to BLEU and ROUGE is built, and it can integrate the studies for different images. This work relates vision to language and has applications for accessibility, vision-based search, and vision understanding

Article Details

Section
Articles