Skip to main content
Google Gemini is a powerful AI model developed by Google, supporting conversational and text generation functions. Currently, ComfyUI has integrated the Google Gemini API, allowing you to directly use the related nodes in ComfyUI to complete conversational functions.

What Google Gemini is good at

  • Conversational and text generation: Chat with Google’s multimodal AI directly inside a ComfyUI workflow
  • Multimodal reasoning: Interpret images with Gemini’s reasoning capabilities
  • Image-to-prompt interpretation: The official template ships with a prompt that turns your images into corresponding drawing prompts
  • Multi-image input: Use Batch Images to send several images for AI interpretation in a single run

Use it in ComfyUI

Google Gemini workflow

Run the Gemini conversational workflow in ComfyUI, locally or on Comfy Cloud