Back to blogs

Multimodal AI / Assistive Tech

See Beyond: AI image captioning for the visually impaired

A vision-language AI project designed to help visually impaired users access visual content through real-time, context-rich captions and auditory feedback.

Sajid Shaikh presenting the See Beyond multimodal AI project in a lecture hall
Presenting See Beyond and discussing the VizWiz dataset, image captioning, and accessibility-focused AI with fellow AI practitioners.

What we built

I had the privilege of presenting the solution we built for See Beyond, a group project that uses AI to enhance the lives of visually impaired users by generating detailed image captions in real time.

The core of the project uses advanced vision-language models such as BLIP-2 and LLaVA, fine-tuned with real-world datasets like VizWiz to generate meaningful, context-rich descriptions from images. We used PyTorch, Hugging Face Transformers, React, and a Flask backend to create a seamless interface with speech-to-text and audio description workflows.

Key technical highlights

Key takeaways

The most important outcome was real-world impact: enabling visually impaired users to access visual content through auditory feedback. Working with VizWiz also forced us to handle real, imperfect images rather than polished benchmark data, which made the data processing and model evaluation work more meaningful.

BLIP-2 outperformed LLaVA in our project setup, showing strong potential for assistive image captioning systems when paired with the right dataset, evaluation strategy, and user-centered interface.

Future prospects

Huge thanks to Professor Xuezhe Ma, TA Nan Xu, Joshua, and my teammates Anika, Yashvi, Nisarg, and Nayan for their collaboration and support.

Discussion

React or ask a follow-up

Comments and reactions are powered by GitHub Discussions under the connectwithsajid brand.