https://amslaurea.unibo.it/id/eprint/25772/
End-to-end Deep Metric Learning con Vision-Language Model per il Fashion Image Captioning - AMS...
language modelfashion imageenddeepmetric
https://liner.com/review/human-attention-in-image-captioning-dataset-and-analysis
Human Attention in Image Captioning: Dataset and Analysis [Quick Review]
Regarding this ICCV 2019 paper, this review summarizes a novel dataset for human attention in image captioning and integrates image saliency for improved m...
image captioningquick reviewhumanattentiondataset
https://www.skills.google/public_profiles/6cdacbc5-fdfd-4f5a-8f45-11f385946772/badges/4434864?locale=es
Create Image Captioning Models | Google Skills
create imagegoogle skillscaptioningmodels
https://research.facebook.com/publications/cross-domain-image-captioning-with-discriminative-finetuning/
Cross-Domain Image Captioning with Discriminative Finetuning - Meta Research
We show that fine-tuning an out-of-the-box neural captioner helps to recover a plain, visually descriptive language that is more informative about image...
cross domainimage captioningfinetuningmetaresearch
https://buescholar.bue.edu.eg/artificial_intelligence/16/
"Arabic Image Captioning: The Effect of Text Pre-processing on the Atte" by Nahla Barakat and Moaz...
Image captioning using deep neural networks has recently gained increasing attention, mostly for English langue, with only few studies in other languages. Good...
image captioningthe effectarabictextpre
https://www.ijisae.org/index.php/IJISAE/article/view/3089
Unveiling the Resilience of Image Captioning Models and the Influence of Pre-trained Models on Deep...
image captioningunveilingresiliencemodelsinfluence
https://www.datacamp.com/tutorial/llama-3-2-90b
Llama 3.2 90B Tutorial: Image Captioning App With Streamlit & Groq | DataCamp
Learn how to build an image captioning app using Streamlit for the front end, Llama 3.2 90B for generating captions, and Groq as the API.
image captioningllamatutorialappstreamlit
https://research.ibm.com/publications/self-critical-sequence-training-for-image-captioning
Self-critical sequence training for image captioning for CVPR 2017 - IBM Research
Self-critical sequence training for image captioning for CVPR 2017 by Steven Rennie et al.
training forimage captioningibm researchselfcritical
https://www.techscience.com/cmc/v75n3/52580/pdf
CMC | Fine-Grained Features for Image Captioning
Image captioning involves two different major modalities (image and sentence) that convert a given image into a language that adheres to visual semantics....
for imagecmcfinefeaturescaptioning
https://oecd.ai/en/catalogue/metric-use-cases/scaling-up-vision-language-pre-training-for-image-captioning
Scaling Up Vision-Language Pre-training for Image Captioning - OECD.AI
The recent advances in neural language models have also been successfully applied to the field of chemistry, offering generative solutions for classical...
scaling uppre trainingfor imagevisionlanguage
https://www.kaggle.com/datasets/aishrules25/automatic-image-captioning-for-visually-impaired
Image Captioning for Visually Impaired people | Kaggle
Data for Visually Impaired to perform Image Captioning task
for visually impairedimage captioningpeoplekaggle
https://www.ijcai.org/proceedings/2017/563
MAT: A Multimodal Attentive Translator for Image Captioning | IJCAI
Electronic proceedings of IJCAI 2017
for imagematmultimodalattentivetranslator
https://aclanthology.org/2017.clicit-1.40/
Deep Learning for Automatic Image Captioning in Poor Training Conditions - ACL Anthology
Caterina Masotti, Danilo Croce, Roberto Basili. Proceedings of the Fourth Italian Conference on Computational Linguistics (CLiC-it 2017). 2017.
deep learningimage captioningautomaticpoortraining
https://cdnjs.deepai.org/publication/unpaired-image-captioning-by-image-level-weakly-supervised-visual-concept-recognition
Unpaired Image Captioning by Image-level Weakly-Supervised Visual Concept Recognition | DeepAI
Mar 7, 2022 - 03/07/22 - The goal of unpaired image captioning (UIC) is to describe images without using image-caption pairs in the training phase. Althoug...
image captioningby levelunpairedsupervisedvisual
https://ijece.iaescore.com/index.php/IJECE/article/view/36630/0
A comprehensive survey on automatic image captioning-deep learning techniques, datasets and...
A comprehensive survey on automatic image captioning-deep learning techniques, datasets and evaluation parameters
image captioningdeep learningcomprehensivesurveyautomatic
https://www.futurebeeai.com/dataset/multi-modal-dataset/french-image-caption-dataset
French Image Captioning Dataset
An image captioning dataset featuring a diverse range of images, each accompanied by multiple captions in French.
image captioningfrenchdataset
https://bytez.com/docs/arxiv/1809.04144/paper
End-to-end Image Captioning Exploits Multimodal Distributional Similarity | Read Paper on Bytez
Sep 11, 2018 - We hypothesize that end-to-end neural image captioning systems work seemingly well because they exploit and learn `distributional similarity' in a multimodal...
image captioningread paperendexploitsmultimodal
https://www.mathworks.com/matlabcentral/fileexchange/75470-image-captioning-app?s_tid=blogs_rc_6
Image Captioning APP - File Exchange - MATLAB Central
Download and share free MATLAB code, including functions, models, apps, support packages and toolboxes
image captioningfile exchangematlab centralapp
https://researchwith.njit.edu/en/publications/a-dual-feature-based-adaptive-shared-transformer-network-for-imag/
A Dual-Feature-Based Adaptive Shared Transformer Network for Image Captioning - New Jersey...
for imagenew jerseydualfeaturebased
https://samim.io/p/2018-03-29-neural-baby-talk-httpsarxivorgabs180309845pytor/
Neural Baby Talk - a novel framework for image captioning that can pro... - samim
samim.io - blogging, research, projects, ideas
baby talkfor imageneuralnovelframework
https://lrec.elra.info/lrec2026-main-744
DEJIMA: A Novel Large-scale Japanese Dataset for Image Captioning and Visual Question Answering -...
a novellarge scalefor imagequestion answeringjapanese
https://isjem.com/download/visionary-ai-multimodal-image-captioning-using-blip-2/
Visionary AI: Multimodal Image Captioning Using Blip-2 - ISJEM Journal
Apr 15, 2026 - Visionary AI: Multimodal Image Captioning Using Blip-2 J. Janaki Ram, M. Siddardha, D. Vinodh Kumar, G. Sunil Kumar, B. Anjanadevi Department of Information...
image captioningvisionaryaimultimodalusing
https://arxiv.org/abs/2111.04193?ref=dataphoenix.info
[2111.04193] Machine-in-the-Loop Rewriting for Creative Image Captioning
Abstract page for arXiv paper 2111.04193: Machine-in-the-Loop Rewriting for Creative Image Captioning
in the loopimage captioningmachinerewritingcreative
https://www.thejournal.club/c/paper/392664/
Transparent Human Evaluation for Image Captioning
for imagetransparenthumanevaluationcaptioning
https://slogix.in/machine-learning/neural-attention-for-image-captioning-review-of-outstanding-methods/
Neural Attention for Image Captioning: Review | S-Logix
In this survey, provide a review of literature related to attentive deep learning models for image captioning.
for imageneuralattentioncaptioningreview
https://research.buaa.edu.cn/en/publications/multiscale-methods-for-optical-remote-sensing-image-captioning/
Multiscale Methods for Optical Remote-Sensing Image Captioning - Beihang University
remote sensingimage captioningbeihang universitymultiscalemethods
https://papers.nips.cc/paper_files/paper/2020/hash/24bea84d52e6a1f8025e313c2ffff50a-Abstract.html
Diverse Image Captioning with Context-Object Split Latent Spaces
image captioningdiversecontextobjectsplit
https://pyimagesearch.com/2025/08/25/meet-blip-the-vision-language-model-powering-image-captioning/
Meet BLIP: The Vision-Language Model Powering Image Captioning - PyImageSearch
Aug 24, 2025 - Discover how BLIP evolved from early captioning models to a powerful vision-language foundation model ready for real-world image captioning deployment.
the visionlanguage modelimage captioningmeetblip
https://mcml.ai/publications/msg+21/
MCML - Scene Graph Generation for Better Image Captioning?
Details on publication MSG+21
scene graphimage captioninggenerationbetter
https://www.aionlinecourse.com/ai-basics/image-captioning
What is Image Captioning | AI Basics | AI Online Course
Artificial intelligence basics: Image Captioning explained! Learn about types, benefits, and factors to consider when choosing an Image Captioning.
what isimage captioningai basicsonline course
https://researchportalplus.anu.edu.au/en/publications/partially-supervised-image-captioning-2/
Partially-supervised image captioning - The Australian National University
image captioningthe australiannational universitysupervised
https://www.runalph.ai/notebooks/huggingface/image-captioning-3
Image Captioning - Hugging Face
Mar 15, 2026 - A notebook by Hugging Face on Alph.
image captioninghugging face
https://wandb.ai/telidavies/ml-news/reports/Crossmodal-3600-Google-s-New-Multilingual-Multicultural-Image-Captioning-Dataset--VmlldzoyNzg5Nzky
Crossmodal-3600: Google's New Multilingual, Multicultural Image Captioning Dataset
image captioninggooglenewmultilingualmulticultural
https://www.apfelpatient.de/en/news/apple-sets-new-standards-in-ai-image-captioning
Apple sets new standards in AI image captioning Apfelpatient
Mar 26, 2026 - Apple surprises with efficient AI: Smaller models deliver better results. Why this could shape the future of AI.
new standardsai imageapplesetscaptioning
https://www.amazon.science/publications/nice-cvpr-2023-challenge-on-zero-shot-image-captioning
NICE: CVPR 2023 challenge on zero-shot image captioning - Amazon Science
In this report, we introduce NICE (New frontiers for zero-shot Image Captioning Evaluation) project1 and share the results and outcomes of 2023 challenge. This...
zero shotimage captioningamazon sciencenicecvpr
https://discovery.researcher.life/article/multimodal-transformer-with-multi-view-visual-representation-for-image-captioning/a367644eac7434158d9f151b2cf3debc
Multimodal Transformer With Multi-View Visual Representation for Image Captioning - R Discovery
Article on Multimodal Transformer With Multi-View Visual Representation for Image Captioning, published in IEEE Transactions on Circuits and Systems for Video...
multi viewfor imagemultimodaltransformervisual