Skip to content

AntResearchNLP/VisualReasoner

 
 

Repository files navigation

VisualReasoner

Official repository for the EMNLP 2024 paper "From the Least to the Most: Building a Plug-and-Play Visual Reasoner via Data Synthesis"


⚙️ Setup

git clone https://github.com/steven-ccq/VisualReasoner.git
cd VisualReasoner

Environment

# Python 3.8
pip install -r requirements.txt

Grounding DINO

cd tools
git clone https://github.com/AntResearchNLP/VisualReasoner.git
cd GroundingDINO/
pip install -e .
mkdir weights
cd weights
wget -q https://github.com/IDEA-Research/GroundingDINO/releases/download/v0.1.0-alpha/groundingdino_swint_ogc.pth
cd ..

Planner Model

Download the adapter.

Merge it with llava-1.5-7b-hf to obtain the Planner model.

Rename the Planner model as planner and move it into models/.

🚀 Inference

First, download the corresponding test sets as guided in the data/ directory.

To facilitate usage, we have provided scripts for each test task:

# TextVQA
bash textvqa.sh
# TallyQA
bash tallyqa.sh
# ST-VQA
bash stvqa.sh
# GQA
bash gqa.sh

The parameters used in the scripts are described in the table below:

Argument Description
input Path to the input file
output Path to the output file
vlm_module Path to the Answer model
src Path to the image folder
model Path to the Planner model
grounding_basedir Path to the Grounding DINO

🎯 Evaluation

# TextVQA
python eval/eval_textvqa.py --input=textvqa.json
# TallyQA
python eval/eval_tallyqa.py --input=tallyqa.json
# ST-VQA
https://rrc.cvc.uab.es/?ch=11
# GQA
python eval/eval_gqa.py --input=gqa.json

🎈 Data

We also provide a 1M dataset synthesized using the least-to-most method. You can access this dataset through 🤗VisualReasoner-1M. We also release a variant of this dataset, which contains 30k end-to-end reasoning processes. You can access this dataset through 🤗VisualReasoner-30k.

@inproceedings{cheng-etal-2024-least,
    title = "From the Least to the Most: Building a Plug-and-Play Visual Reasoner via Data Synthesis",
    author = "Cheng, Chuanqi  and
      Guan, Jian  and
      Wu, Wei  and
      Yan, Rui",
    editor = "Al-Onaizan, Yaser  and
      Bansal, Mohit  and
      Chen, Yun-Nung",
    booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.emnlp-main.284/",
    doi = "10.18653/v1/2024.emnlp-main.284",
    pages = "4941--4957",
    abstract = "We explore multi-step reasoning in vision-language models (VLMs). The problem is challenging, as reasoning data consisting of multiple steps of visual and language processing are barely available. To overcome the challenge, we first introduce a least-to-most visual reasoning paradigm, which interleaves steps of decomposing a question into sub-questions and invoking external tools for resolving sub-questions. Based on the paradigm, we further propose a novel data synthesis approach that can automatically create questions and multi-step reasoning paths for an image in a bottom-up manner. Our approach divides the complex synthesis task into a few simple sub-tasks, and (almost entirely) relies on open-sourced models to accomplish the sub-tasks. Therefore, the entire synthesis process is reproducible and cost-efficient, and the synthesized data is quality guaranteed. With the approach, we construct 50k visual reasoning examples. Then, we develop a visual reasoner through supervised fine-tuning, which is capable of generally enhancing the reasoning abilities of a wide range of existing VLMs in a plug-and-play fashion. Extensive experiments indicate that the visual reasoner can consistently and significantly improve four VLMs on four VQA benchmarks."
}

About

[EMNLP 2024] Official repository for paper "From the Least to the Most: Building a Plug-and-Play Visual Reasoner via Data Synthesis"

Resources

Stars

Watchers

Forks

Releases

Packages

Contributors

Languages