Google Gemini

This ComfyUI workflow leverages Google's Gemini model to demonstrate the capabilities of multimodal AI in generating coherent and contextually relevant text. The workflow utilizes several nodes, including the GeminiNode, which serves as the core processing unit, and the LoadImage node, which allows users to input images that the Gemini model can analyze and interpret. By integrating these components, the workflow showcases how Gemini's advanced reasoning capabilities can be applied to both text and image inputs, providing a seamless experience for generating text based on visual content. Additionally, the PreviewAny and BatchImagesNode nodes facilitate the visualization and management of multiple outputs, making it easier to handle and review generated content.

API

Use this workflow from code

Every Comfy workflow is a JSON graph. The payload below is this workflow, exactly as ComfyUI runs it — fetch it from the URL, keep it in version control, or load it in ComfyUI and run it node by node.

api_google_gemini.json
Fetching workflow JSON…

Run it from TypeScript or Python with the Comfy SDK. The same code targets Comfy Cloud or a ComfyUI you host yourself — only the base URL changes.

// Install (beta)
npm i @comfyorg/sdk

// Run "Google Gemini" (TypeScript)
import { Comfy } from "@comfyorg/sdk";

const client = new Comfy({ apiKey: "comfyui-..." });

// This workflow, exported in API format (see note below)
const wf = await client.workflows.fromFile("api_google_gemini_api.json");

const asset = client.assets.fromFile("input.png");
wf.setInput("2", "image", asset); // LoadImage

const job = await client.run(wf);
await job.getOutputs("20")[0].toFile("output.png"); // SaveText

The SDK takes a workflow in API format: open this workflow in ComfyUI and use File → Export Workflow (API).

Comfy Cloud API access requires a plan with an API key.

FAQ

Frequently Asked Questions

View all workflows
Showing 30 of 30 templates