Two routes to multimodal understanding
Compare OCR plus BLIP-2 captioning with a CLIP-based approach to harmful-content detection.
Inside the project
Research showcaseHow it works.
Read the image and text
OCR extracts written content while BLIP-2 adds image-caption information.
Continue in the source.
Open the source notebook in Jupyter, Colab, or the environment described in the README. Data and model downloads may be required.
git clone https://github.com/eforus-overseer/Harmfull-Content-Detection-Classification-BLIP2-OCR-CLIP.gitRead the setup and requirements ↗Project artifacts.
Source links point to the original public repository. Credit belongs to the project authors and the dependencies credited there.