Analogist: Out-of-the-box Visual In-Context Learning with Image Diffusion Model

There is no author summary for this article yet. Authors can add summaries to their articles on ScienceOpen to make them more accessible to a non-specialist audience.

Abstract

Visual In-Context Learning (ICL) has emerged as a promising research area due to its capability to accomplish various tasks with limited example pairs through analogical reasoning. However, training-based visual ICL has limitations in its ability to generalize to unseen tasks and requires the collection of a diverse task dataset. On the other hand, existing methods in the inference-based visual ICL category solely rely on textual prompts, which fail to capture fine-grained contextual information from given examples and can be time-consuming when converting from images to text prompts. To address these challenges, we propose Analogist, a novel inference-based visual ICL approach that exploits both visual and textual prompting techniques using a text-to-image diffusion model pretrained for image inpainting. For visual prompting, we propose a self-attention cloning (SAC) method to guide the fine-grained structural-level analogy between image examples. For textual prompting, we leverage GPT-4V's visual reasoning capability to efficiently generate text prompts and introduce a cross-attention masking (CAM) operation to enhance the accuracy of semantic-level analogy guided by text prompts. Our method is out-of-the-box and does not require fine-tuning or optimization. It is also generic and flexible, enabling a wide range of visual tasks to be performed in an in-context manner. Extensive experiments demonstrate the superiority of our method over existing approaches, both qualitatively and quantitatively.

Related collections

Author and article information

Journal

Publication date Created: 16 May 2024

Article

ArXiV ID: 2405.10316

SO-VID: 5da90148-c670-426d-9fb1-e3411046f313

License:

http://creativecommons.org/licenses/by/4.0/

History

Custom metadata

Comments Project page: https://analogist2d.github.io

Categories cs.CV cs.GR

ScienceOpen disciplines: Computer vision & Pattern recognition,Graphics & Multimedia design

Data availability:

ScienceOpen disciplines: Computer vision & Pattern recognition, Graphics & Multimedia design

Analogist: Out-of-the-box Visual In-Context Learning with Image Diffusion Model

Read this article at

Abstract

Related collections

Recursive Rule based Visual Categorization

Author and article information

Journal

Article

History

Custom metadata

Comments

Comment on this article

Similar content 52