Generate text captions for images from their CLIP embeddings. Includes PyTorch model code and example training script.

Frank Odom

Last update: Dec 21, 2022

Related tags

Deep Learning clip-text-decoder

Overview

clip-text-decoder

Generate text captions for images from their CLIP embeddings. Includes PyTorch model code and example training script.

Example Predictions

Example captions were computed with the pretrained model mentioned below.

"A man riding a wave on top of a surfboard."

A baseball player is swinging a bat at a ball.

"A dog running across a field with a frisbee."

Installation

Install for easier access to the following objects/classes:

clip_text_decoder.datasets.ClipCocoCaptionsDataset
clip_text_decoder.models.ClipDecoder
clip_text_decoder.models.ClipDecoderInferenceModel
clip_text_decoder.tokenizer.Tokenizer

The train.py script will not be available in the installed package, since it's located in the root directory. To train new models, either clone this repository or recreate train.py locally.

Using pip:

pip install clip-text-decoder

From source:

git clone https://github.com/fkodom/clip-text-decoder.git
cd clip-text-decoder
pip install .

NOTE: You'll also need to install openai/CLIP to encode images with CLIP. This is also required by ClipCocoCaptionsDataset to build the captions dataset the first time (cached for subsequent calls).

pip install "clip @ git+https://github.com/openai/CLIP.git"

For technical reasons, the CLIP dependency can't be included in the PyPI package, since it's not an officially published package.

Training

Launch your own training session using the provided script (train.py):

python train.py --max-epochs 5

Training CLI arguments, along with their default values:

--max-epochs 5  # (int)
--num-layers 6  # (int)
--dim-feedforward 256  # (int)
--precision 16  # (16 or 32)
--seed 0  # (int)

Inference

The training script will produce a model.zip archive, containing the Tokenizer and trained model parameters. To perform inference with it:

import clip
from PIL import Image
import torch

from clip_text_decoder.model import ClipDecoderInferenceModel

device = "cuda" if torch.cuda.is_available() else "cpu"
model = ClipDecoderInferenceModel.load("path/to/model.zip").to(device)
clip_model, clip_preprocessor = clip.load("ViT-B/32", device=device, jit=False)

# Create a blank dummy image
dummy_image = Image.new("RGB", (224, 224))
preprocessed = clip_preprocessor(dummy_image).to(device)
# Add a batch dimension using '.unsqueeze(0)'
encoded = clip_model.encode_image(preprocessed.unsqueeze(0))
text = model(encoded)

print(text)
# Probably some nonsense, because we used a dummy image.

Pretrained Models

A pretrained CLIP decoder is hosted in my Google Drive, and can easily be downloaded by:

from clip_text_decoder.model import ClipDecoderInferenceModel

model = ClipDecoderInferenceModel.download_pretrained()

To cache the pretrained model locally, so that it's not re-downloaded each time:

model = ClipDecoderInferenceModel.download_pretrained("/path/to/model.zip")

Shortcomings

Only works well with COCO-style images. If you go outside the distribution of COCO objects, you'll get nonsense text captions.
Relatively short training time. Even within the COCO domain, you'll occasionally see incorrect captions. Quite a few captions will have bad grammar, repetitive descriptors, etc.

Comments

Decoding Text Embeddings Coded Using Hugging Face ClipTextModel

Suppose that I have text embeddings created using Hugging Face's ClipTextModel using the following method:

import torch
from transformers import CLIPTokenizer, CLIPTextModel

class_list = ["i love going home and playing with my wife and kids", "i love going home", "playing with my wife and kids", 
"family", "war", "writing"]

model = CLIPTextModel.from_pretrained("openai/clip-vit-large-patch14")
tokenizer = CLIPTokenizer.from_pretrained("openai/clip-vit-large-patch14")

inputs = tokenizer(class_list, padding=True, return_tensors="pt")
outputs = model(**inputs)
hidden_state = outputs.last_hidden_state
embeddings = outputs.pooler_output

Questions:

Is It possible to use the clip-text-decoder to convert the embeddings back to text?
If it is indeed possible to do so, could you provide an example of how?

Looking forward to receiving your feedback.

opened by mbdzi 6

Fix string error when loading clip models.

error

The model name string ( VIT-xxx ) in the check_vision_backbone function is not compatible with the model name string ( ViT-xxx ) of the clip repository, which will cause at least one error in check_vision_backbone function or when loading the clip model.

solution

In this PR, the model name string in the check_vision_backbone function is modified to ViT-xxx to make it compatible with the clip repository.

opened by Adenialzz 1
BLIP vision backbone
Added blip backbone; still cleaning up last pieces

Bug fixes for training script, and remove debug code.

Fix dependencies in test workflow; update README statistics

Fix test issue with CUDA device

Update unit tests for newer Python, torch versions

Test up to Python 3.10

Test up to Python 3.9

Install lavis first
opened by fkodom 0
Feature: Beam Search
Add beam search, clip dependency to setup.py

Fix installation instructions

Remove main clause

Add '--beam-size' option to 'train.py' script.

Update README; propagate the '--beam-size' arg through eval functions

Update setup.cfg, add pre-commit hooks

Reformat images

Remove fixed image width

Add detail to README; comments to call method for beam search

Updated README headline
opened by fkodom 0
Bug Fixes for Broken Tests
Cache the old fashioned way :)

Fix silly typo in test for image caption model

Apply black and isort formatting

Install latest version of 'black', reapply formatting

Fix flake8 issue (duplicate function definition), and install latest patch version of pytorch for tests.

Skip slow tests by default, add 'slow' marker to inference model tests.
opened by fkodom 0
GPT2 Decoder
Update model to use DistilGPT2 as a pre-trained decoder.

Removed tokenizer (no longer used), fixed bugs in Model source file, and updated model unit tests.

Backwards compatibility for 'gdown.download' method.

Update installation requirements, caption examples in README
opened by fkodom 0
Upgrade CodeSee workflow to version 2
CodeSee is a code visibility platform.

This change updates the CodeSee workflow file to the latest version for security, maintenance, and support improvements (see changelog below).

That workflow file:

runs CodeSee's code analysis on every PR push and merge

uploads that analysis to CodeSee.

It does not transmit your code.

The code analysis is used to generate maps and insights about this codebase.

CodeSee workflow changelog:

Improved security: Updates permission to be read-only.

Improved future maintenance: Replaces the body of the workflow with a single github action: codesee-action. This makes it significantly easier for CodeSee to introduce future improvements and fixes without requiring another PR like this.

Improved Python support: The action now properly supports Python 3.11, and will continue to support new Python versions as they are released.
opened by codesee-maps[bot] 1

Incompatible checksum error

I see the following error when trying to load the pretrained model.

    tokenizer=pickle.loads(tokenizer_buffer.read()),
  File "stringsource", line 6, in spacy.pipeline.trainable_pipe.__pyx_unpickle_TrainablePipe
_pickle.PickleError: Incompatible checksums (102742709 vs 0x417ddeb = (cfg, model, name, vocab))

Am I missing something?

opened by dapurv5 0

Releases(1.4.4)

1.4.4(Nov 7, 2022)
What's Changed

Fix string error when loading clip models. by @Adenialzz in https://github.com/fkodom/clip-text-decoder/pull/12

New Contributors

@Adenialzz made their first contribution in https://github.com/fkodom/clip-text-decoder/pull/12

Full Changelog: https://github.com/fkodom/clip-text-decoder/compare/1.4.3...1.4.4
Source code(tar.gz)
Source code(zip)
1.4.3(Nov 7, 2022)
What's Changed

Refactor Dataset by @fkodom in https://github.com/fkodom/clip-text-decoder/pull/11

Full Changelog: https://github.com/fkodom/clip-text-decoder/compare/1.4.2...1.4.3
Source code(tar.gz)
Source code(zip)
1.4.2(Oct 26, 2022)
What's Changed

Huggingface Evaluate by @fkodom in https://github.com/fkodom/clip-text-decoder/pull/9

Full Changelog: https://github.com/fkodom/clip-text-decoder/compare/1.4.1...1.4.2
Source code(tar.gz)
Source code(zip)
1.4.1(Oct 26, 2022)
What's Changed

Datapipes by @fkodom in https://github.com/fkodom/clip-text-decoder/pull/8

Full Changelog: https://github.com/fkodom/clip-text-decoder/compare/1.4.0...1.4.1
Source code(tar.gz)
Source code(zip)
1.4.0(Oct 23, 2022)
What's Changed

BLIP vision backbone by @fkodom in https://github.com/fkodom/clip-text-decoder/pull/7

Full Changelog: https://github.com/fkodom/clip-text-decoder/compare/1.3.0...1.4.0
Source code(tar.gz)
Source code(zip)
1.3.0(Oct 2, 2022)
What's Changed

Feature: Beam Search by @fkodom in https://github.com/fkodom/clip-text-decoder/pull/5

Bug Fix: PyPI Release by @fkodom in https://github.com/fkodom/clip-text-decoder/pull/6

Full Changelog: https://github.com/fkodom/clip-text-decoder/compare/1.2.0...1.3.0
Source code(tar.gz)
Source code(zip)
1.2.0(Jan 29, 2022)
What's Changed

Cache CLIP embeddings for the dataset, rather than recomputing them each time.

Reduce model file sizes by storing at lower precision

Add an ImageCaptionInferenceModel class for easier out-of-the-box use

Fix some broken unit tests

Better Data Caching by @fkodom in https://github.com/fkodom/clip-text-decoder/pull/3

Bug Fixes for Broken Tests by @fkodom in https://github.com/fkodom/clip-text-decoder/pull/4

Full Changelog: https://github.com/fkodom/clip-text-decoder/compare/1.1.0...1.2.0
Source code(tar.gz)
Source code(zip)
1.1.0(Dec 22, 2021)
What's Changed

GPT2 Decoder by @fkodom in https://github.com/fkodom/clip-text-decoder/pull/2

New Contributors

@fkodom made their first contribution in https://github.com/fkodom/clip-text-decoder/pull/2

Full Changelog: https://github.com/fkodom/clip-text-decoder/compare/1.0.0...1.1.0
Source code(tar.gz)
Source code(zip)
1.0.0(Nov 15, 2021)

Source code(tar.gz)
Source code(zip)
0.1.1(Nov 14, 2021)

Add installation docs to README, and automatically publish to PyPI using GitHub Actions workflow.
Source code(tar.gz)
Source code(zip)
0.1.0(Nov 14, 2021)

First pre-release with pretrained models in README.
Source code(tar.gz)
Source code(zip)

Owner

Frank Odom

Director of Innovation at Plainsight. I like neural nets, and neural nets like me.

GitHub

Python package to generate image embeddings with CLIP without PyTorch/TensorFlow

imgbeddings A Python package to generate embedding vectors from images, using OpenAI's robust CLIP model via Hugging Face transformers. These image em

81 Jan 4, 2023

This is the code for our KILT leaderboard submission to the T-REx and zsRE tasks. It includes code for training a DPR model then continuing training with RAG.

KGI (Knowledge Graph Induction) for slot filling This is the code for our KILT leaderboard submission to the T-REx and zsRE tasks. It includes code fo

72 Jan 6, 2023

Video-Captioning - A machine Learning project to generate captions for video frames indicating the relationship between the objects in the video

1 Jan 23, 2022

Rename Images with Auto Generated Neural Image Captions

Recaption Images with Generated Neural Image Caption Example Usage: Commandline: Recaption all images from folder /home/feng/Downloads/images to folde

3 May 1, 2022

An image base contains 490 images for learning (400 cars and 90 boats), and another 21 images for testingAn image base contains 490 images for learning (400 cars and 90 boats), and another 21 images for testing

SVM Données Une base d’images contient 490 images pour l’apprentissage (400 voitures et 90 bateaux), et encore 21 images pour fait des tests. Prétrait

3 Nov 30, 2021

In this project we investigate the performance of the SetCon model on realistic video footage. Therefore, we implemented the model in PyTorch and tested the model on two example videos.

Contrastive Learning of Object Representations Supervisor: Prof. Dr. Gemma Roig Institutions: Goethe University CVAI - Computational Vision & Artifici

6 Dec 8, 2022

Source code for "MusCaps: Generating Captions for Music Audio" (IJCNN 2021)

MusCaps: Generating Captions for Music Audio Ilaria Manco1 2, Emmanouil Benetos1, Elio Quinton2, Gyorgy Fazekas1 1 Queen Mary University of London, 2

57 Dec 7, 2022

This YoloV5 based model is fit to detect people and different types of land vehicles, and displaying their density on a fitted map, according to their coordinates and detected labels.

This YoloV5 based model is fit to detect people and different types of land vehicles, and displaying their density on a fitted map, according to their

8 May 22, 2022

A simple editor for captions in .SRT file extension

WaySRT A simple editor for captions in .SRT file extension The program doesn't use any external dependecies, just run: python way_srt.py {file_name.sr

3 Nov 16, 2022

This is the official source code for SLATE. We provide the code for the model, the training code, and a dataset loader for the 3D Shapes dataset. This code is implemented in Pytorch.

SLATE This is the official source code for SLATE. We provide the code for the model, the training code and a dataset loader for the 3D Shapes dataset.

66 Dec 26, 2022

A PyTorch Lightning solution to training OpenAI's CLIP from scratch.

train-CLIP ?? A PyTorch Lightning solution to training CLIP from scratch. Goal ⚽ Our aim is to create an easy to use Lightning implementation of OpenA

396 Dec 30, 2022

Source code for models described in the paper "AudioCLIP: Extending CLIP to Image, Text and Audio" (https://arxiv.org/abs/2106.13043)

AudioCLIP Extending CLIP to Image, Text and Audio This repository contains implementation of the models described in the paper arXiv:2106.13043. This

458 Jan 2, 2023

Neon-erc20-example - Example of creating SPL token and wrapping it with ERC20 interface in Neon EVM

Example of wrapping SPL token by ERC2-20 interface in Neon Requirements Install

7 Mar 28, 2022

Simple implementation of OpenAI CLIP model in PyTorch.

It was in January of 2021 that OpenAI announced two new models: DALL-E and CLIP, both multi-modality models connecting texts and images in some way. In this article we are going to implement CLIP model from scratch in PyTorch. OpenAI has open-sourced some of the code relating to CLIP model but I found it intimidating and it was far from something short and simple. I also came across a good tutorial inspired by CLIP model on Keras code examples and I translated some parts of it into PyTorch to build this tutorial totally with our beloved PyTorch!

226 Jan 5, 2023

Softlearning is a reinforcement learning framework for training maximum entropy policies in continuous domains. Includes the official implementation of the Soft Actor-Critic algorithm.

Softlearning Softlearning is a deep reinforcement learning toolbox for training maximum entropy policies in continuous domains. The implementation is

997 Dec 30, 2022

Generate text captions for images from their CLIP embeddings. Includes PyTorch model code and example training script.

Related tags

Overview

clip-text-decoder

Example Predictions

Installation

Training

Inference

Pretrained Models

Shortcomings

Comments

Releases(1.4.4)

1.4.4(Nov 7, 2022)

What's Changed

New Contributors

1.4.3(Nov 7, 2022)

What's Changed

1.4.2(Oct 26, 2022)

What's Changed

1.4.1(Oct 26, 2022)

What's Changed

1.4.0(Oct 23, 2022)

What's Changed

1.3.0(Oct 2, 2022)

What's Changed

1.2.0(Jan 29, 2022)

What's Changed

1.1.0(Dec 22, 2021)

What's Changed

New Contributors

1.0.0(Nov 15, 2021)

0.1.1(Nov 14, 2021)

0.1.0(Nov 14, 2021)

Owner

Frank Odom

Python package to generate image embeddings with CLIP without PyTorch/TensorFlow

This is the code for our KILT leaderboard submission to the T-REx and zsRE tasks. It includes code for training a DPR model then continuing training with RAG.

Video-Captioning - A machine Learning project to generate captions for video frames indicating the relationship between the objects in the video

Rename Images with Auto Generated Neural Image Captions

An image base contains 490 images for learning (400 cars and 90 boats), and another 21 images for testingAn image base contains 490 images for learning (400 cars and 90 boats), and another 21 images for testing

In this project we investigate the performance of the SetCon model on realistic video footage. Therefore, we implemented the model in PyTorch and tested the model on two example videos.

Source code for "MusCaps: Generating Captions for Music Audio" (IJCNN 2021)

This YoloV5 based model is fit to detect people and different types of land vehicles, and displaying their density on a fitted map, according to their coordinates and detected labels.

A simple editor for captions in .SRT file extension

This is the official source code for SLATE. We provide the code for the model, the training code, and a dataset loader for the 3D Shapes dataset. This code is implemented in Pytorch.

A PyTorch Lightning solution to training OpenAI's CLIP from scratch.

Source code for models described in the paper "AudioCLIP: Extending CLIP to Image, Text and Audio" (https://arxiv.org/abs/2106.13043)

Neon-erc20-example - Example of creating SPL token and wrapping it with ERC20 interface in Neon EVM

Simple implementation of OpenAI CLIP model in PyTorch.

Softlearning is a reinforcement learning framework for training maximum entropy policies in continuous domains. Includes the official implementation of the Soft Actor-Critic algorithm.

TAP: Text-Aware Pre-training for Text-VQA and Text-Caption, CVPR 2021 (Oral)

Character-Input - Create a program that asks the user to enter their name and their age

Generating images from caption and vice versa via CLIP-Guided Generative Latent Space Search

Example-custom-ml-block-keras - Custom Keras ML block example for Edge Impulse